A method, device, equipment and storage medium for evaluating spoken pronunciation
By using an acoustic feature recognition model with the diphone state as the modeling unit in oral pronunciation evaluation, the problem of insufficient acoustic modeling capabilities of the HMM-DNN model is solved, and the accuracy and effect of the evaluation are improved.
Patent Information
- Application Number
- CN202110150522.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-03
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-02-03
AI Technical Summary
The existing HMM-DNN acoustic model has weak acoustic modeling capabilities in oral pronunciation evaluation, resulting in poor speech recognition performance and low accuracy of evaluation results.
The acoustic feature recognition model with the diphone state as the modeling unit is used for oral pronunciation evaluation, and the posterior probability of the phoneme is determined through the acoustic likelihood probability vector, and then the pronunciation results are evaluated.
It improves the accuracy and effectiveness of oral pronunciation evaluation, and has better acoustic modeling and speech recognition capabilities than the HMM-DNN model.
Smart Images

Figure CN113571094B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a spoken pronunciation evaluation method, device, equipment and storage medium. Background Art
[0002] Nowadays, learning knowledge and skills through educational applications (Application, APP) has become a common learning method for users. In a common application scenario, an APP used to help users learn foreign languages can provide an oral pronunciation practice function, which can evaluate and score the user's oral pronunciation based on the pronunciation audio uploaded by the user, so that the user can understand whether his or her oral pronunciation is standard.
[0003] In the related technology, the Hidden Markov Model-Deep Neural Networks (HMM-DNN) acoustic model is currently mainly used to evaluate the user's spoken pronunciation. The HMM-DNN acoustic model uses triphones as modeling units, determines and outputs the acoustic posterior probability based on the acoustic features of the pronunciation audio uploaded by the user; then, the user's spoken pronunciation evaluation result is determined by a binary classifier based on the acoustic posterior probability output by the HMM-DNN model.
[0004] However, in practical applications, the acoustic modeling capability of the above-mentioned HMM-DNN acoustic model is relatively weak, and the speech recognition performance is relatively poor. When evaluating the user's spoken pronunciation based on the acoustic posterior probability output by the HMM-DNN acoustic model, the accuracy of the evaluation results is relatively low, and the evaluation effect is often not ideal. Summary of the invention
[0005] The embodiments of the present application provide a method, apparatus, device and storage medium for evaluating spoken pronunciation, which can ensure that the determined spoken pronunciation evaluation results have high accuracy and effectively improve the effect of spoken pronunciation evaluation.
[0006] In view of this, the first aspect of the present application provides a method for evaluating spoken pronunciation, the method comprising:
[0007] Acquire a target audio to be evaluated; the target audio corresponds to a target text;
[0008] Performing acoustic feature extraction processing on the target audio to obtain a target acoustic feature sequence;
[0009] Determining an acoustic likelihood probability vector according to the target acoustic feature sequence through an acoustic feature recognition model; the acoustic feature recognition model is a model that uses a diphone state as a modeling unit;
[0010] Determining a posterior probability of a target phoneme in the target text based on the acoustic likelihood probability vector and the target text;
[0011] The target pronunciation evaluation result is determined according to the posterior probability of the target phoneme.
[0012] A second aspect of the present application provides a spoken pronunciation evaluation device, the device comprising:
[0013] An audio acquisition module, used to acquire a target audio to be evaluated; the target audio corresponds to a target text;
[0014] An acoustic feature extraction module, used to perform acoustic feature extraction processing on the target audio to obtain a target acoustic feature sequence;
[0015] A likelihood probability determination module, used to determine an acoustic likelihood probability vector according to the target acoustic feature sequence through an acoustic feature recognition model; the acoustic feature recognition model is a model using a diphone state as a modeling unit;
[0016] A posterior probability determination module, used to determine the posterior probability of the target phoneme in the target text based on the acoustic likelihood probability vector and the target text;
[0017] The pronunciation evaluation module is used to determine the target pronunciation evaluation result according to the posterior probability of the target phoneme.
[0018] A third aspect of the present application provides a device, the device comprising a processor and a memory:
[0019] The memory is used to store computer programs;
[0020] The processor is used to execute the steps of the spoken pronunciation evaluation method as described in the first aspect according to the computer program.
[0021] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the steps of the oral pronunciation evaluation method described in the first aspect.
[0022] In a fifth aspect, the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps of the oral pronunciation evaluation method described in the first aspect above.
[0023] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0024] The embodiment of the present application provides a method for evaluating spoken pronunciation, which innovatively uses an acoustic feature recognition model with a two-phoneme state as a modeling unit to evaluate spoken pronunciation. Through the acoustic feature recognition model, an acoustic likelihood probability vector is determined according to a target acoustic feature sequence corresponding to the target audio to be evaluated; then, based on the acoustic likelihood probability vector and the target text, the posterior probability of the target phoneme in the target text is determined; finally, the target pronunciation evaluation result is determined according to the posterior probability of the target phoneme. Considering that the acoustic feature recognition model with two-phone states as modeling units has better acoustic modeling ability and speech recognition ability than the HMM-DNN acoustic model in the related technology, the embodiment of the present application introduces the acoustic feature recognition model into the oral pronunciation evaluation process; and in order to make the acoustic likelihood probability vector output by the acoustic feature recognition model suitable for oral pronunciation evaluation, the embodiment of the present application also proposes an implementation method for determining the acoustic posterior probability based on the acoustic likelihood probability; in this way, the acoustic feature recognition model with two-phone states as modeling units is used for oral pronunciation evaluation, which can ensure that the oral pronunciation evaluation results with higher accuracy are obtained, thereby effectively improving the oral pronunciation evaluation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A schematic diagram of an application scenario of the spoken pronunciation evaluation method provided in an embodiment of the present application;
[0026] Figure 2 A flowchart of a method for evaluating spoken pronunciation provided in an embodiment of the present application;
[0027] Figure 3 A schematic diagram of the HMM topology used by the Chain model provided in the embodiment of the present application;
[0028] Figure 4 A schematic diagram of an exemplary HMM topology structure provided in an embodiment of the present application;
[0029] Figure 5 A schematic diagram of an exemplary pronunciation evaluation result display interface provided in an embodiment of the present application;
[0030] Figure 6 A flowchart of another method for evaluating spoken pronunciation provided in an embodiment of the present application;
[0031] Figure 7 A schematic diagram of the structure of a spoken pronunciation evaluation device provided in an embodiment of the present application;
[0032] Figure 8A schematic diagram of the structure of another oral pronunciation evaluation device provided in an embodiment of the present application;
[0033] Fig. 9 A schematic diagram of the structure of another oral pronunciation evaluation device provided in an embodiment of the present application;
[0034] Fig.10 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application;
[0035] Fig.11 A schematic diagram of the structure of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0036] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0037] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0038] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0039] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0040] The key technologies of speech technology include: Automatic Speech Recognition (ASR), Text to Speech (TTS) and voiceprint recognition. Enabling computers to listen, see, speak and feel is the future development direction of human-computer interaction, among which speech has become one of the most promising human-computer interaction methods in the future.
[0041] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, drones, robots, smart medical care, smart customer service, etc. I believe that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0042] The solution provided in the embodiments of the present application relates to artificial intelligence speech technology, which is specifically described by the following embodiments:
[0043] In the related art, the HMM-DNN acoustic model with triphones as modeling units is usually used to evaluate spoken pronunciation. However, the acoustic modeling ability and speech recognition ability of the HMM-DNN acoustic model are relatively weak. Accordingly, the pronunciation evaluation results determined based on the acoustic posterior probability output by the HMM-DNN acoustic model are often less accurate, and the pronunciation evaluation results obtained are poor.
[0044] In response to the problems existing in the above-mentioned related technologies, an embodiment of the present application provides a method for evaluating spoken pronunciation. The method innovatively uses an acoustic feature recognition model with two-phoneme states as modeling units for spoken pronunciation evaluation, which can ensure that the determined pronunciation evaluation results have high accuracy and achieve better pronunciation evaluation effects.
[0045] Specifically, in the oral pronunciation evaluation method provided in the embodiment of the present application, the target audio to be evaluated is first obtained, and the target audio corresponds to the target text. Then, the target audio is subjected to acoustic feature extraction processing to obtain a target acoustic feature sequence. Next, an acoustic likelihood probability vector is determined based on the target acoustic feature sequence through an acoustic feature recognition model; the acoustic feature recognition model here is a model with a two-phoneme state as a modeling unit. Furthermore, based on the acoustic likelihood probability vector and the target text, the posterior probability of the target phoneme in the target text is determined. Finally, based on the posterior probability of the target phoneme, the target pronunciation evaluation result is determined.
[0046] Since the acoustic feature recognition model with two-phone state as the modeling unit has better acoustic modeling ability and speech recognition ability than the HMM-DNN model with three-phoneme as the modeling unit, the method provided in the embodiment of the present application introduces the acoustic feature recognition model into the oral pronunciation evaluation process; and in order to make the acoustic likelihood probability vector output by the acoustic feature recognition model suitable for oral pronunciation evaluation, the embodiment of the present application also proposes an implementation method for determining the acoustic posterior probability based on the acoustic likelihood probability. In this way, the acoustic feature recognition model with two-phone state as the modeling unit is used for oral pronunciation evaluation, and the use of the acoustic feature recognition model for oral pronunciation evaluation can ensure that the determined pronunciation evaluation result has a high degree of accuracy, thereby effectively improving the oral pronunciation evaluation effect.
[0047] It should be understood that the oral pronunciation evaluation method provided in the embodiment of the present application can be applied to devices with speech processing capabilities, such as terminal devices, servers, etc. The terminal device can specifically be a smart phone, a computer, a tablet computer, a personal digital assistant (PDA), an intelligent speaker, an intelligent robot, etc. The server can specifically be an application server or a Web server, and in actual deployment, it can be a stand-alone server, a cluster server, or a cloud server.
[0048] In order to facilitate understanding of the spoken pronunciation evaluation method provided in the embodiment of the present application, the application scenario of the spoken pronunciation evaluation method is first exemplarily introduced below.
[0049] See also Figure 1 , Figure 1 Schematic diagram of the application scenario of the oral pronunciation evaluation method provided in the embodiment of the present application. Figure 1 As shown, the application scenario includes a terminal device 110 and a server 120, and the terminal device 110 and the server 120 can communicate through a network. Among them, the terminal device 110 is installed with a target application, and the target application has a spoken pronunciation practice function; the server 120 is used to execute the spoken pronunciation evaluation method provided in the embodiment of the present application.
[0050] In actual applications, the user can use the oral pronunciation practice function in the target application through the terminal device 110. For example, when the user uses the oral pronunciation practice function, the terminal device 110 can display the follow-up text provided by the target application to the user, and can collect the audio generated when the user reads the follow-up text in response to the user's operation on the audio recording control; after the terminal device 110 detects that the user confirms that the audio recording is completed, it can send the collected audio to the server 120 through the network.
[0051] After receiving the audio sent by the terminal device 110, the server 120 may regard the audio as the target audio to be evaluated, and regard the follow-up text based on which the user records the audio as the target text. The server 120 performs acoustic feature extraction processing on the target audio to obtain a corresponding target acoustic feature sequence, which includes the acoustic features in each time unit in the target audio.
[0052] Then, the server 120 can determine the acoustic likelihood probability vector according to the above target acoustic feature sequence through the acoustic feature recognition model, where the acoustic feature recognition model is a model that uses the diphone state as a modeling unit. Exemplarily, the above acoustic feature recognition model can be a chain model, which can output a corresponding acoustic likelihood probability vector according to the input acoustic feature sequence. The acoustic likelihood probability vector is essentially a T*N dimensional matrix, where T represents the number of time units included in the acoustic feature sequence, N is the total number of diphone states, and the element X in the acoustic likelihood probability vector is ij represents the likelihood probability corresponding to the acoustic feature in the i-th time unit under the j-th diphone state, which represents the conditional probability of observing the acoustic feature in the i-th time unit given the j-th diphone state.
[0053] Next, the server 120 can determine the posterior probability of the target phoneme in the target text based on the acoustic likelihood probability vector output by the acoustic feature recognition model and the target text. Specifically, the server 120 can forcibly align the target acoustic feature sequence corresponding to the target audio with the target text, that is, determine the acoustic features in the target acoustic feature sequence corresponding to the target two phonemes in the target text, and the time interval to which the acoustic features belong is the target time interval corresponding to the target two phonemes. The server 120 can subsequently evaluate whether the user's pronunciation of the target two phonemes is standard based on the acoustic features in the target time interval.
[0054] Since the acoustic likelihood probability vector output by the acoustic feature recognition model is usually difficult to be used directly for evaluating pronunciation, the server 120 needs to determine the acoustic posterior probability that can be used for evaluating pronunciation based on the acoustic likelihood probability vector. The posterior probability refers to the probability of observing a certain phoneme state under the condition of given acoustic features. In specific implementation, the server 120 can determine the posterior probability of the target phoneme based on the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector; the target phoneme here can be the target two phonemes themselves, or the latter of the target two phonemes.
[0055] Furthermore, the server 120 may determine the target pronunciation evaluation result based on the posterior probability of the target phoneme. For example, the server 120 may evaluate whether the user's pronunciation of a target phoneme is accurate based on the posterior probability of the target phoneme. For another example, the server 120 may also evaluate whether the user's pronunciation of a word is accurate based on the posterior probabilities of multiple target phonemes belonging to a certain word. For another example, the server 120 may also evaluate whether the user's pronunciation of a sentence is accurate based on the posterior probabilities of multiple target phonemes belonging to a sentence; and so on.
[0056] After the server 120 determines the target pronunciation evaluation result through the above process, it can send the target pronunciation evaluation result to the terminal device 110 through the network, so that the terminal device 110 can display the target pronunciation evaluation result to the user, so that the user can understand whether his or her spoken pronunciation is standard.
[0057] It should be understood that Figure 1 The application scenarios shown are only examples. In practical applications, the acoustic feature recognition model can also be deployed locally on the terminal device 110, and the terminal device 110 independently performs spoken pronunciation evaluation based on the target audio input by the user. The application scenarios of the spoken pronunciation evaluation method provided in the embodiment of the present application are not limited in any way.
[0058] The oral pronunciation evaluation method provided by the present application is described in detail below through a method embodiment.
[0059] See also Figure 2 , Figure 2 The following is a flow chart of the oral pronunciation evaluation method provided in the embodiment of the present application. For the convenience of description, the following embodiment is introduced by taking the execution subject of the oral pronunciation evaluation method as a server as an example. Figure 2 As shown, the spoken pronunciation evaluation method comprises the following steps:
[0060] Step 201: Acquire target audio to be evaluated; the target audio corresponds to a target text.
[0061] In actual applications, the server may obtain the audio of the spoken pronunciation to be evaluated as the target audio, and the text corresponding to the target audio as the target text.
[0062] In a possible implementation, the server can obtain the audio sent by the terminal device as the target audio to be evaluated. Exemplarily, a target application with an oral pronunciation practice function is installed in the terminal device, and the oral pronunciation practice function can provide the user with a follow-up text, and display the follow-up text in the interface corresponding to the oral pronunciation practice function; the user can trigger the terminal device to collect the audio generated when the user reads the follow-up text by touching the start follow-up control, and trigger the terminal device to stop collecting audio by touching the end follow-up control; the terminal device sends the collected audio to the server through the network, so that the server uses the received audio as the target audio to be evaluated, and the target text corresponding to the target audio is the user's follow-up text.
[0063] In addition, the above-mentioned oral pronunciation practice function can also support users to play freely, that is, in the absence of a text to follow, the user can trigger the terminal device to collect the audio generated by his free reading by touching the start follow-up control, and trigger the terminal device to stop collecting audio by touching the end follow-up control; the terminal device sends the collected audio to the server through the network, so that the server uses the received audio as the target audio to be evaluated. At this time, the server can determine the target text corresponding to the target audio by performing speech recognition on the target audio.
[0064] It should be understood that the above-mentioned implementation method of users triggering the terminal device to collect audio is only an example. In actual applications, users can also trigger the terminal device to collect audio in other ways. For example, they can trigger the terminal device to collect audio by long pressing the audio entry control. This application does not impose any limitation on the implementation method of triggering the terminal device to collect audio.
[0065] In another possible implementation, the server can obtain the target audio to be evaluated from the database and determine the target text corresponding to the target audio. For example, the audio uploaded by the user can be stored in the database first, and when a pronunciation evaluation is required for a certain audio, the server can retrieve the audio from the database as the target audio and determine the target text corresponding to the target audio.
[0066] It should be understood that in practical applications, the server may also obtain the target audio to be evaluated and the target text corresponding to the target audio by other means, and the present application does not limit the method of obtaining the target audio and the target text. Moreover, the target audio to be evaluated in the embodiment of the present application is not limited to the audio input by the user, but may also be other types of audio.
[0067] Step 202: Perform acoustic feature extraction processing on the target audio to obtain a target acoustic feature sequence.
[0068] After the server obtains the target audio, it can perform acoustic feature extraction processing on the target audio to obtain a target acoustic feature sequence corresponding to the target audio.
[0069] Exemplarily, the server can obtain the target acoustic feature sequence corresponding to the target audio by performing a series of processes such as pre-emphasis, frame windowing, decoding, discrete Fourier transform, Mel filtering, logarithm, discrete cosine transform and differential extraction on the target audio. Of course, in practical applications, the server can also perform acoustic feature extraction processing on the target audio in other ways, and this application does not make any limitation on the implementation method of extracting the target acoustic feature sequence from the target audio.
[0070] It should be noted that the target acoustic feature sequence includes acoustic features within multiple time units, and these multiple time units correspond to the duration of the target audio. That is, the target audio is divided into multiple sub-audio segments from the time dimension, and the acoustic features corresponding to a sub-audio segment are the acoustic features within a time unit, and the acoustic features corresponding to each sub-audio segment constitute the above-mentioned target acoustic feature sequence. The length of the above-mentioned time unit can be set according to actual needs, such as 1ms, 10ms, etc., and this application does not impose any limitation on the length of the time unit.
[0071] Step 203: determining an acoustic likelihood probability vector according to the target acoustic feature sequence through an acoustic feature recognition model; the acoustic feature recognition model is a model that uses a diphone state as a modeling unit.
[0072] After the server performs acoustic feature extraction processing on the target audio and obtains the target acoustic feature sequence, it can call the acoustic feature recognition model and input the target acoustic feature sequence into the acoustic feature recognition model. The acoustic feature recognition model analyzes and processes the input target acoustic feature sequence and outputs an acoustic likelihood probability vector accordingly.
[0073] It should be noted that the above-mentioned acoustic feature recognition model is a neural network model with diphone states (hereinafter referred to as senone states) as modeling units, which is used to determine the likelihood probability corresponding to each acoustic feature in each time unit in the acoustic feature sequence under each senone state for the input acoustic feature sequence. The likelihood probability specifically refers to the conditional probability of observing the acoustic feature under a given senone state.
[0074] For example, in the method provided in the embodiment of the present application, the acoustic feature recognition model can be specifically a Chain model. Different from the HMM-DNN acoustic model that uses triphones as modeling units, the Chain model is trained based on the sequence discrimination training criterion and uses diphone states as modeling units. The Chain model usually uses a two-state HMM topology to represent a diphone. Figure 3 The figure shows the schematic diagram of the HMM topology used by the Chain model, where the diphone p 1 p 2 The corresponding first senone state a can only appear once, while the diphone p 1 p 2 The corresponding second senone state b can appear any number of times, that is, it can appear zero times, one time, or multiple times. The number of occurrences of the second senone state b depends on the diphone p 1 p 2 In other words, for a diphone, once the time interval length of its corresponding acoustic feature is determined, its corresponding senone state sequence can also be determined accordingly. For example, assuming a diphone p 1 p 2 The time interval length of the corresponding acoustic feature is three time units, then the two phonemes p 1 p 2 The corresponding senone state sequence should be [abb].
[0075] In addition, the output results of the Chain model are also different from those of the HMM-DNN acoustic model. The output results of the HMM-DNN acoustic model are the acoustic posterior probability, that is, the conditional probability of observing the senone state under a given acoustic feature, while the output results of the Chain model are the acoustic likelihood probability, that is, the conditional probability of observing the acoustic feature under a given senone state.
[0076] Assume that a diphone p 1 p 2 The alignment time is from the 1st time unit to the Tth time unit. The acoustic feature sequence O in this time interval is [o 1 o 2 …o T ]∈R F×T , F represents the dimension of acoustic features; assuming that the senone state sequence corresponding to this time interval is S = [s 1 s 2 …s T ],s t Represents the senone state at time t. The Chain model is used to determine the given senone state s tObserve below t The conditional probability P θ (o t |s t ). According to the independence assumption, given a senone state sequence S in a certain time interval, the Chain model can calculate the likelihood probability through formula (1):
[0077]
[0078] Among them, θ represents the parameters of the Chain model, and the right side of formula (1) can be regarded as the likelihood function with respect to θ.
[0079] It should be understood that in practical applications, in addition to using the Chain model as an acoustic feature recognition model, other neural network models that use diphone states as modeling units and are used to determine acoustic likelihood probabilities can also be used as acoustic feature recognition models in the embodiments of the present application. No limitation is made to the acoustic feature recognition models in the embodiments of the present application.
[0080] It should be noted that the acoustic likelihood probability vector output by the acoustic feature recognition model can be understood as a T*N-dimensional likelihood probability matrix, where T represents the number of time units included in the input target acoustic feature sequence, and N represents the total number of senone states. In the acoustic likelihood probability vector, the element X ij It represents the likelihood probability corresponding to the acoustic feature in the i-th time unit under the j-th senone state, that is, the conditional probability of observing the acoustic feature in the i-th time unit given the j-th senone state.
[0081] Step 204: Determine the posterior probability of the target phoneme in the target text based on the acoustic likelihood probability vector and the target text.
[0082] In specific implementation, the server can determine the time interval to which the acoustic features corresponding to the target two phonemes in the target text in the target acoustic feature sequence belong based on the acoustic likelihood probability vector and the target text, as the target time interval; and then determine the posterior probability of the target phoneme based on the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector.
[0083] That is, after the server obtains the acoustic likelihood probability vector output by the acoustic feature recognition model, it can force the target acoustic feature sequence corresponding to the target audio and the target text corresponding to the target audio to be aligned based on the acoustic likelihood probability vector. That is, for the target two phonemes in the target text (which can be any two phonemes in the target text), determine the time interval to which the corresponding acoustic features in the target acoustic feature sequence belong as the target time interval.
[0084] In one possible implementation, the server may construct a candidate senone state sequence corresponding to the target text based on the duration of the target audio and the senone states corresponding to each diphone in the target text; then, for each candidate senone state sequence, based on the acoustic likelihood probability vector output by the acoustic feature recognition model, determine the reference likelihood probability corresponding to the candidate senone state sequence; then, based on the reference likelihood probability corresponding to each candidate senone state sequence, select a target senone state sequence from each candidate senone state sequence; finally, for each diphone in the target text, based on the target senone state sequence, determine the time interval to which the acoustic feature corresponding to the diphone in the target acoustic feature sequence belongs.
[0085] Specifically, given the duration of the target audio, the server can allocate a corresponding time interval for each diphone in the target text, and construct a senone state sequence corresponding to the diphone according to the length of the time interval corresponding to each diphone and the senone state corresponding to the diphone. According to the arrangement order of each diphone in the target text, the senone state sequences corresponding to each diphone are connected in series to obtain a candidate senone state sequence corresponding to the target text. The server adjusts the time interval allocated to each diphone in the target text, and repeats the above operation to obtain multiple candidate senone state sequences corresponding to the target text.
[0086] Then, the server can determine the corresponding reference likelihood probability for each candidate senone state sequence. Specifically, the server can determine the corresponding time unit for each senone state in the candidate senone state sequence, and find the likelihood probability corresponding to the acoustic feature in the time unit under the senone state in the acoustic likelihood probability vector output by the acoustic feature recognition model as the likelihood probability corresponding to the senone state; then, based on the likelihood probabilities corresponding to each senone state in the candidate senone state sequence, the reference likelihood probability corresponding to the candidate senone state sequence is calculated. For example, the sum or product of the likelihood probabilities corresponding to each senone state in the candidate senone state sequence can be calculated as the reference likelihood probability corresponding to the candidate senone state sequence.
[0087] Furthermore, the server can select the best candidate senone state sequence from each candidate senone state sequence as the target senone state sequence according to the reference likelihood probability corresponding to each candidate senone state sequence. For example, the server can select the candidate senone state sequence with the largest corresponding reference likelihood probability as the target senone state sequence.
[0088] Each senone state included in the target senone state sequence corresponds one-to-one to each time unit in the target audio, and the target senone state sequence is composed of the senone state sequence corresponding to each diphone in the target text; based on this, the server can, for each diphone in the target text, connect in series the time units corresponding to the senone states in the senone state sequence corresponding to the diphone to obtain the time interval corresponding to the diphone.
[0089] To facilitate understanding of the above implementation process, the following example takes the target text as hi, and allows silence before hi (the corresponding phoneme marker is eps), and the duration of the target audio includes 5 time units. Figure 4 The HMM topology structure corresponding to hi shown is used to illustrate the above implementation process.
[0090] The target text hi includes the monophones sil, h, i, and sil. Considering that silence eps is allowed before hi, the target text hi includes the following diphones: (eps, sil), (sil, h), (h, i), (i, sil); the HMM topology corresponding to the target text hi is as follows Figure 4 As shown, the senone states corresponding to the two phonemes (eps, sil) include a1 and b1, the senone states corresponding to the two phonemes (sil, h) include a2 and b2, the two phonemes (h, i) include a3 and b3, and the senone states corresponding to the two phonemes (i, sil) include a4 and b4.
[0091] When the duration of the target audio includes 5 time units, the server can allocate the first to third time units to the diphones (eps, sil), (sil, h) and (h, i) respectively, and allocate the fourth and fifth time units to the diphone (i, sil). In this case, the candidate senone state sequence constructed by the server is a1a2a3a4b4; the server can also construct other candidate senone state sequences, such as a1a2a3b3a4, a1a2b2a3a4, and a1b1a2a3a4, by adjusting the allocation method of the time intervals of each diphone.
[0092] Then, the server can determine the corresponding reference likelihood probability for each candidate senone state sequence; taking the determination of the corresponding reference likelihood probability for the candidate senone state sequence a1a2a3b3a4 as an example, the server can find the likelihood probability P(o) corresponding to the acoustic feature in the first time unit under senone state a1 in the acoustic likelihood probability vector output by the acoustic feature recognition model. 1 |s 1 =a 1 ), the likelihood probability P(o) corresponding to the acoustic feature in the second time unit under senone state a2 2 |s 2 =a 2 ), the likelihood probability P(o) corresponding to the acoustic feature in the third time unit under senone state a3 3 |s 3 =a 3 ), the likelihood probability P(o) corresponding to the acoustic feature in the fourth time unit under senone state a4 4 |s 4 =a 4 ), and the likelihood probability P(o) corresponding to the acoustic feature in the fifth time unit under senone state b4 5 |s 5 =b 4 ); Then, based on the above likelihood probability, the reference likelihood probability corresponding to the candidate senone state sequence a1a2a3b3a4 is calculated as P(o 1 |s 1 =a 1 )×P(o 2 |s 2 =a 2 )×P(o 3 |s 3 =a 3 )×P(o 4 |s 4 =a 4 )×P(o 5 |s 5 =b 4 ). Thus, in a similar manner, the corresponding reference likelihood probabilities are calculated for the candidate senone state sequences a1a2a3b3a4, a1a2b2a3a4, and a1b1a2a3a4, respectively.
[0093] Furthermore, the server can determine the candidate senone state sequence corresponding to the maximum reference likelihood probability as the target senone state sequence based on the reference likelihood probabilities corresponding to each candidate senone state sequence; and according to the target senone state sequence, determine the corresponding time intervals for each diphone (eps, sil), (sil, h), (h, i), (i, sil) in the target text hi. Assuming that the target senone state sequence is a1a2a3b3a4, it can be determined that the first time unit of the target audio corresponds to the diphone (eps, sil), the second time unit of the target audio corresponds to the diphone (sil, h), the third time unit of the target audio corresponds to the diphone (h, i), and the fourth and fifth time units of the target audio correspond to the diphone (i, sil).
[0094] The server forcibly aligns the target acoustic feature sequence corresponding to the target audio with the target text corresponding to the target audio, and after determining the target time interval corresponding to the target two phonemes in the target text, it can further determine the posterior probability of the target phoneme based on the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector.
[0095] It should be noted that the above-mentioned target phoneme can refer to the target diphone itself, or the latter monophone in the target diphone (the reason is that the latter monophone in the diphone is the main phone of the diphone). The posterior probability of the above-mentioned target phoneme can refer to an independent posterior probability value, or a posterior probability distribution composed of multiple posterior probability values. When the target phoneme is the target diphone, the posterior probability distribution of the target phoneme is an M*M dimensional posterior probability distribution, where M is the number of all monophones, and the element Y in the posterior probability distribution is ij Represents the posterior probability corresponding to the diphone composed of the i-th monophone and the j-th monophone; when the target phoneme is the latter monophone in the target diphone, the posterior probability distribution of the target phoneme is an M*1-dimensional posterior probability distribution, where M is the number of all monophones, and the element Z in the posterior probability distribution is i1 Represents the posterior probability corresponding to the i-th monophone.
[0096] When specifically determining the posterior probability of the target phoneme, the server can determine the reference HMM topology based on the length of the target time interval corresponding to the target two phonemes; then, combine the monophones in pairs to obtain multiple candidate two phonemes corresponding to the acoustic features in the target time interval; further, determine the senone state sequence corresponding to each candidate two phoneme based on the senone state corresponding to each candidate two phoneme and the reference HMM topology; finally, based on each senone state sequence, determine the posterior probability of the target phoneme based on the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector output by the acoustic feature recognition model.
[0097] After the server determines the target time interval, it can determine the reference HMM topology of the two-phone applicable to the target time interval according to the length of the target time interval; for example, when the target time interval only includes one time unit, the reference HMM topology only includes the main senone state corresponding to the two-phone; when the target time interval includes two time units, the reference HMM topology includes both the main senone state and the slave senone state corresponding to the two-phone, and the slave senone state only appears once; when the target time interval includes three time units, the reference HMM topology includes both the main senone state and the slave senone state corresponding to the two-phone, and the slave senone state appears twice in a cycle; and so on.
[0098] In addition, the server needs to combine each monophone in pairs to obtain multiple candidate diphones; for example, according to the rules of the CMU pronunciation dictionary, without considering the position and stress, a total of 39 monophones are involved. At this time, combining each monophone in pairs will obtain 39*39=1521 candidate diphones. Then, for each candidate diphone, the server can determine the senone state sequence corresponding to the candidate diphone according to the senone state corresponding to the candidate diphone and the above reference HMM topology; for example, assuming that the diphone p 1 p 2 The corresponding senone states include the master senone state a and the slave senone state b. In the case where the reference HMM topology corresponds to three time units, for the diphone p 1 p 2 The constructed senone state sequence should be abb.
[0099] Furthermore, the server can determine the posterior probability of the target phoneme based on each senone state sequence and the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector output by the acoustic feature recognition model. The present application embodiment provides four exemplary implementation methods for determining the posterior probability of the target phoneme, and these four implementation methods are introduced below.
[0100] In a first possible implementation, the server may determine the posterior probability distribution of the diphone as the posterior probability of the target phoneme in the following manner: for each senone state sequence, determine the reference likelihood probability of the candidate diphone corresponding to the senone state sequence according to the likelihood probability corresponding to the acoustic features in the target time interval under the senone state included in the senone state sequence in the acoustic likelihood probability vector; determine the sum of the reference likelihood probabilities of the candidate diphones as the total reference likelihood probability; for each candidate diphone, determine the posterior probability of the candidate diphone according to the reference likelihood probability of the candidate diphone and the total reference likelihood probability; and further, construct the posterior probability distribution of the diphone corresponding to the acoustic features of the target time interval based on the posterior probabilities of the candidate diphones as the posterior probability of the target phoneme.
[0101] Specifically, the server can search the acoustic likelihood probability vector for each senone state in the senone state sequence for the likelihood probability corresponding to the acoustic feature in the time unit corresponding to the senone state, as the likelihood probability corresponding to the senone state; then, according to the likelihood probability corresponding to each senone state in the senone state sequence, determine the reference likelihood probability of the candidate diphone corresponding to the senone state sequence, for example, the sum or product of the likelihood probabilities corresponding to each senone state in the senone state sequence can be calculated as the reference likelihood probability of the candidate diphone corresponding to the senone state sequence. Then, the sum of the reference likelihood probabilities of each candidate diphone is calculated as the total reference likelihood probability. For each candidate diphone, the ratio between the reference likelihood probability of the candidate diphone and the total reference likelihood probability is calculated as the posterior probability of the candidate diphone. Then, the posterior probability of each candidate diphone is used to construct the diphone posterior probability distribution as the posterior probability of the target phoneme. For example, assuming there are M monophones, M monophones are combined two by two to obtain M*M candidate diphones. The posterior probability of each of the M*M candidate diphones can be used to construct an M*M dimensional diphone posterior probability distribution, where the element Y ij is the posterior probability of a candidate biphone consisting of the i-th monophone and the j-th monophone.
[0102] Assume that a candidate diphone is p 1 p 2 , the candidate diphone p 1 p 2 The reference likelihood probability is P(O|p 1 p 2 ), then the candidate diphone p can be calculated by formula (2) 1 p 2 The posterior probability P(p1 p 2 |O):
[0103]
[0104] Among them, q 1 q 2 Can represent any candidate diphone, P(O|q 1 q 2 ) represents a candidate diphone q 1 q 2 The reference likelihood probability of the candidate diphone p 1 p 2 When the corresponding senone state sequence is S, P(O|p 1 p 2 )=P(O|S).
[0105] In this way, the posterior probability of each candidate diphone is calculated by formula (2), and the posterior probability of each candidate diphone can be used to construct the posterior probability distribution of the diphone corresponding to the target time interval, and use it as the posterior probability of the target phoneme.
[0106] In a second possible implementation, the server can determine the posterior probability value of the target diphone itself as the posterior probability of the target phoneme in the following manner: for each senone state sequence, determine the reference likelihood probability of the candidate diphone corresponding to the senone state sequence according to the likelihood probability corresponding to the acoustic features in the target time interval under the senone state included in the senone state sequence in the acoustic likelihood probability vector; determine the sum of the reference likelihood probabilities of each candidate diphone as the total reference likelihood probability; for the target diphone, determine the posterior probability of the target diphone as the posterior probability of the target phoneme according to the reference likelihood probability of the target diphone and the total reference likelihood probability.
[0107] Different from the first implementation method mentioned above, in the second implementation method, the server can determine the posterior probability of the target diphone only for the target diphone in the target text, and directly use the posterior probability of the target diphone as the posterior probability of the target phone. That is to say, in the second implementation method, the server also needs to determine the reference likelihood probability for each candidate diphone, and determine the total reference likelihood probability based on the reference likelihood probability of each candidate diphone; however, the server does not need to calculate the posterior probability for each candidate diphone, but only needs to calculate the posterior probability for the target diphone, that is, it only needs to calculate the ratio of the reference likelihood probability of the target diphone to the total reference likelihood probability, and obtain the posterior probability of the target diphone as the posterior probability of the target phone.
[0108] For example, assuming the target diphone is p 1p 2 , the target two phonemes p 1 p 2 The reference likelihood probability is P(O|p 1 p 2 ), then the target two-phoneme p can be calculated by formula (3): 1 p 2 The posterior probability P(p 1 p 2 |O):
[0109]
[0110] Among them, q 1 q 2 Can represent any candidate diphone, P(O|q 1 q 2 ) represents a candidate diphone q 1 q 2 The reference likelihood probability of the target two phonemes p 1 p 2 When the corresponding senone state sequence is S, P(O|p 1 p 2 )=P(O|S).
[0111] In this way, the posterior probability of the target two-phoneme is calculated by formula (3), and the posterior probability of the target two-phoneme can be used as the posterior probability of the target phoneme.
[0112] It should be noted that in the first implementation and the second implementation, when the server determines the posterior probability of a diphone (a candidate diphone or a target diphone), the prior probability of the diphone may also be comprehensively considered. Taking the calculation of the posterior probability of the target diphone as an example, the server may determine the posterior probability of the target diphone based on the reference likelihood probability of the target diphone, the prior probability of the target diphone, the total reference likelihood probability, and the prior probability of each candidate diphone.
[0113] For example, assuming the target diphone is p 1 p 2 , the target two phonemes p 1 p 2 The reference likelihood probability is P(O|p 1 p 2 ), then the target two-phoneme p can be calculated by formula (4): 1 p 2 The posterior probability P(p 1 p 2 |O):
[0114]
[0115] Among them, P(p 1 p 2 ) represents the target diphone p 1 p 2 The prior probability of the target two phonemes p 1 p 2 The number of occurrences in the historical text is determined; P(q 1 q 2 ) represents any candidate diphone q 1 q 2 The prior probability of the candidate diphone q can be obtained by counting 1 q 2 The number of occurrences in the historical text is determined.
[0116] The posterior probability calculation formula shown in the above formula (4) is derived based on the Bayesian formula. The posterior probability calculation formulas shown in the above formulas (2) and (3) are obtained by converting formula (4) under the assumption of equal prior probability. Experimental studies have found that the posterior probabilities calculated by the above formulas (2) and (3) are often more accurate.
[0117] In a third possible implementation, the server may determine the posterior probability distribution of a single phone as the posterior probability of the target phoneme in the following manner: the front single phoneme and the back single phoneme in the candidate two phonemes are regarded as the front phoneme and the back phoneme, respectively; for each senone state sequence corresponding to the candidate two phonemes including the same back phoneme, the reference likelihood probability of the candidate two phonemes is determined according to the likelihood probability corresponding to the acoustic features in the target time interval under the senone state included in the senone state sequence in the acoustic likelihood probability vector; and from the reference likelihood probabilities of each candidate two phonemes including the back phoneme, the maximum reference likelihood probability is selected as the reference likelihood probability of the back phoneme; then, the sum of the reference likelihood probabilities of each back phoneme is determined as the total reference likelihood probability; further, for each back phoneme, the posterior probability of the back phoneme is determined according to the reference likelihood probability of the back phoneme and the total reference likelihood probability; finally, based on the posterior probabilities of each back phoneme, the posterior probability distribution of single phones corresponding to the acoustic features in the target time interval is constructed as the posterior probability of the target phoneme.
[0118] Specifically, for each phoneme, the server can construct a set of candidate diphones corresponding to the post-phoneme using the candidate diphones that take it as the post-phoneme. Then, for each post-phoneme, the reference likelihood probability of the post-phoneme is determined according to the senone state sequence corresponding to each candidate diphone in the corresponding candidate diphone set. In specific implementation, the server can search for the likelihood probability corresponding to the acoustic features in the time unit corresponding to the senone state in the senone state in the acoustic likelihood probability vector for each senone state in the senone state sequence as the likelihood probability corresponding to the senone state; then, according to the likelihood probability corresponding to each senone state in the senone state sequence, the reference likelihood probability of the candidate diphone corresponding to the senone state sequence is determined. For example, the sum or product of the likelihood probabilities corresponding to each senone state in the senone state sequence can be calculated as the reference likelihood probability of the candidate diphone corresponding to the senone state sequence. After determining the reference likelihood probabilities for each candidate diphone in the candidate diphone set corresponding to a postphone, the server may select the maximum reference likelihood probability from the reference likelihood probabilities of each candidate diphone as the reference likelihood probability of the postphone.
[0119] Then, the sum of the reference likelihood probabilities of each postphone is calculated as the total reference likelihood probability. For each postphone, the ratio of the reference likelihood probability of the postphone to the total reference likelihood probability is calculated as the posterior probability of the postphone. Then, the posterior probability distribution of the single phoneme is constructed using the posterior probability of each postphone as the posterior probability of the target phoneme; for example, assuming that there are M single phones in total, the posterior probabilities of the M single phones can be used to construct an M*1-dimensional single phoneme posterior probability distribution, where the element Z i1 is the posterior probability of the ith monophone.
[0120] For example, according to the Bayesian formula, a single phoneme p 2 The posterior probability P(p 2 |O) can be calculated by formula (5):
[0121]
[0122] in, In the monophone p 2 In the case of fixed, including each single phoneme p 1 The sum of the posterior probabilities of the candidate two phones; Represents the sum of the reference likelihood probabilities of each candidate biphone.
[0123] However, experimental studies have shown that the posterior probability of a single phoneme calculated by formula (5) is often not accurate enough; therefore, the calculation formula for the posterior probability of a single phoneme shown in formula (5) is adjusted to formula (6):
[0124]
[0125] in, Indicates the postphoneme p 2 The reference likelihood probability, that is, after fixing the phoneme p 2 , enumerate the front phoneme p 1 In the case of 1 The candidate two phonemes calculate their reference likelihood probabilities, and then include the latter phoneme p 2 The largest reference likelihood probability is selected from the reference likelihood probabilities of each candidate two phonemes as the latter phoneme p 2 The reference likelihood probability. Indicates each post-phoneme q 2 The sum of the respective reference likelihood probabilities.
[0126] The reason for such adjustment is that, in actual applications, not all diphones obtained by pairwise combination of each monophone exist. Some monophones correspond to many diphones, while some monophones correspond to few diphones. Direct summation calculation will result in a larger posterior probability for monophones corresponding to more diphones, and a smaller posterior probability for monophones corresponding to fewer diphones. In practice, it is found that among the reference likelihood probabilities of each candidate diphone including the same postphone, selecting the largest reference likelihood probability to represent the reference likelihood probability of the postphone can make the posterior probability calculated subsequently more accurate.
[0127] In this way, after calculating the posterior probability of each postphone by formula (6), the posterior probability of each postphone can be used to construct the posterior probability distribution of the single phoneme corresponding to the target time interval, and use it as the posterior probability of the target phoneme.
[0128] In a fourth possible implementation, the server may determine the posterior probability value of the target post-phoneme in the target diphone as the posterior probability of the target phoneme in the following manner: the front monophone and the back monophone in the candidate diphone are regarded as the front phoneme and the back phoneme, respectively; for each senone state sequence corresponding to the candidate diphone including the same back phoneme, the reference likelihood probability of the candidate diphone is determined according to the likelihood probability corresponding to the acoustic features in the target time interval under the senone state included in the senone state sequence in the acoustic likelihood probability vector; and from the reference likelihood probabilities of each candidate diphone including the back phoneme, the maximum reference likelihood probability is selected as the reference likelihood probability of the back phoneme; then, the sum of the reference likelihood probabilities of each back phoneme is determined as the total reference likelihood probability; and further, for the target back phoneme in the target diphone, the posterior probability of the target back phoneme is determined according to the reference likelihood probability of the target back phoneme and the total reference likelihood probability as the posterior probability of the target phoneme.
[0129] Different from the third implementation method mentioned above, in the fourth implementation method, the server can determine the posterior probability only for the postphone included in the target two phonemes in the target text (i.e., the target postphone), and directly use the posterior probability of the target postphone as the posterior probability of the target phoneme. That is to say, in the fourth implementation method, the server also needs to determine the reference likelihood probability for each postphone, and determine the total reference likelihood probability based on the reference likelihood probability of each postphone; however, the server does not need to calculate the posterior probability for each postphone, but only needs to calculate the posterior probability for the target postphone included in the target two phonemes, that is, it only needs to calculate the ratio of the reference likelihood probability of the target postphone to the total reference likelihood probability, and obtain the posterior probability of the target postphone as the posterior probability of the target phoneme.
[0130] For example, assuming the target diphone is p 1 p 2 , the target postphoneme is p 2 The server can calculate the target postphoneme p by formula (7) 2 The posterior probability of:
[0131]
[0132] in, Indicates the target postphoneme p 2 The reference likelihood probability of q 1 q 2 can represent any candidate diphone, Indicates a postphoneme q 2 The reference likelihood probability.
[0133] In this way, the posterior probability of the target post-phoneme in the target two phonemes is calculated by formula (7), and the posterior probability of the target post-phoneme can be used as the posterior probability of the target phoneme.
[0134] Step 205: Determine a target pronunciation evaluation result according to the posterior probability of the target phoneme.
[0135] After the server determines the posterior probability of the target phoneme, the target pronunciation evaluation result can be determined according to the posterior probability of the target phoneme.
[0136] In practical applications, the server usually needs to use a pronunciation evaluation model to determine the target pronunciation evaluation result based on the posterior probability of the target phoneme. The pronunciation evaluation model is a neural network model pre-trained in a supervised training manner, that is, the server can use a large number of training samples including the posterior probability of phonemes and annotated pronunciation evaluation results to train the pronunciation evaluation model until the pronunciation evaluation model meets the training end conditions, for example, until the performance of the pronunciation evaluation model reaches a preset performance standard, or until the number of iterative training of the pronunciation evaluation model reaches a preset number of training times, and so on.
[0137] It should be understood that in actual applications, the processing object of the pronunciation evaluation model trained by the server can be any one of the two-phone posterior probability distribution, the two-phone posterior probability value, the monophone posterior probability distribution and the monophone posterior probability value. The processing object of the pronunciation evaluation model can be set according to actual needs. This application does not make any limitation on the processing object of the pronunciation evaluation model.
[0138] It should be noted that, in actual applications, the server can determine the above-mentioned target pronunciation evaluation results in at least one of the following ways: through a phoneme evaluation model, determine the phoneme pronunciation evaluation results according to the posterior probability of the target phoneme; through a word evaluation model, determine the word pronunciation evaluation results according to a first posterior probability set, where the first posterior probability set includes: the posterior probability of each target phoneme included in the word to be evaluated in the target text; through a sentence evaluation model, determine the sentence pronunciation evaluation results according to a second posterior probability set, where the second posterior probability set includes: the posterior probability of each target phoneme included in the sentence to be evaluated in the target text.
[0139] That is, after the server determines the posterior probability of the target phoneme, at least one of the phoneme pronunciation, word pronunciation and sentence pronunciation can be evaluated based on the posterior probability of the target phoneme. When the server evaluates the phoneme pronunciation, the posterior probability of the target phoneme determined by step 205 can be directly input into the phoneme evaluation model, and the result output by the phoneme evaluation model is obtained as the phoneme pronunciation evaluation result. When the server evaluates the word pronunciation, the target phonemes included in the word to be evaluated can be determined, and then the posterior probability of each target phoneme in the word to be evaluated is input into the word evaluation model, and the output result of the word evaluation model is obtained as the word pronunciation evaluation result. When the server evaluates the sentence pronunciation, the target phonemes included in the sentence to be evaluated can be determined, and then the posterior probability of each target phoneme in the sentence to be evaluated is input into the sentence evaluation model, and the output result of the sentence evaluation model is obtained as the sentence pronunciation evaluation result. Of course, the server can further utilize the article evaluation model to determine the article pronunciation evaluation result according to the posterior probability of each target phoneme included in each sentence in the article to be evaluated, and so on.
[0140] It should be understood that in a scenario where the server performs pronunciation evaluation based on the target audio uploaded by the terminal device, after the server determines the target pronunciation evaluation result, it can further return the target pronunciation evaluation result to the terminal device so that the terminal device can display the pronunciation evaluation result to the user. Figure 5 FIG. 1 is a schematic diagram of an exemplary pronunciation evaluation result display interface. Figure 5 As shown, the terminal device can display the sentence pronunciation evaluation result in the interface, and the sentence pronunciation evaluation result can be specifically represented by a score or a star rating; in addition, Figure 5 As shown in (a), the user can click on a phoneme to view the pronunciation evaluation result of the phoneme, such as Figure 5 As shown in (b), the user can view the pronunciation evaluation results of the word by long pressing the word.
[0141] The oral pronunciation evaluation method provided in the embodiment of the present application takes into account that the acoustic feature recognition model with a two-phone state as a modeling unit has better acoustic modeling ability and speech recognition ability than the HMM-DNN model with a three-phoneme as a modeling unit. Therefore, the acoustic feature recognition model is introduced into the oral pronunciation evaluation process; and in order to make the acoustic likelihood probability vector output by the acoustic feature recognition model suitable for oral pronunciation evaluation, the embodiment of the present application also proposes an implementation method for determining the acoustic posterior probability based on the acoustic likelihood probability. In this way, the acoustic feature recognition model with a two-phone state as a modeling unit is used for oral pronunciation evaluation, and the use of the acoustic feature recognition model for oral pronunciation evaluation can ensure that the determined pronunciation evaluation result has a high degree of accuracy, thereby effectively improving the oral pronunciation evaluation effect.
[0142] In order to further understand the oral pronunciation evaluation method provided in the embodiment of the present application, Figure 6 The flowchart shown, taking the determination of pronunciation evaluation results based on the posterior probability distribution of a single phoneme as an example, provides an overall exemplary introduction to the oral pronunciation evaluation method provided in the embodiment of the present application.
[0143] like Figure 6 As shown, the terminal device can send the audio recorded when the user follows the target text to the server through the network, so that the server uses it as the target audio to be evaluated. After the server obtains the target audio, it can first perform acoustic feature extraction processing on the target audio in step 601 to obtain a target acoustic feature sequence corresponding to the target audio.
[0144] Then, the server may determine an acoustic likelihood probability vector based on the target acoustic feature sequence using the Chain model in step 602. The acoustic likelihood probability vector includes the likelihood probabilities corresponding to the acoustic features in each time unit in the target acoustic feature sequence in each senone state.
[0145] Next, the server may perform step 603 to forcibly align the target acoustic feature sequence with the target text according to the acoustic likelihood probability vector output by the Chain model, and determine, for each diphone included in the target text, the time interval to which the corresponding acoustic feature in the target acoustic feature sequence belongs.
[0146] Furthermore, the server can, through steps 604 to 606, take the time interval corresponding to each two-phoneme in the target text as a processing unit, and determine the posterior probability distribution of the monophone corresponding to the acoustic feature in each time interval in the acoustic likelihood probability vector based on the likelihood probability corresponding to the acoustic feature in the time interval.
[0147] In the specific implementation, for the single phoneme p 2 Each phoneme p can be enumerated 1 As the front phoneme, with the single phoneme p 2 Then, for each candidate diphone, the reference likelihood probability of the candidate diphone is determined according to the senone state sequence and acoustic likelihood probability vector corresponding to the candidate diphone; and then, from the postphone p 2 Among the reference likelihood probabilities of each candidate two phonemes, select the largest reference likelihood probability As the postphoneme p 2reference likelihood probability; thus, the reference likelihood probability of each monophone is determined in the above manner. Then, the server can calculate the sum of the reference likelihood probabilities of each monophone as the total reference likelihood probability. Furthermore, for each monophone, the ratio of its reference likelihood probability to the total reference likelihood probability is calculated as the posterior probability of the monophone. Finally, the posterior probability distribution of the monophone corresponding to the acoustic features in the time interval is constructed using the posterior probability of each monophone.
[0148] Finally, the server can use the pre-trained pronunciation evaluation model through step 607 to evaluate and score the spoken pronunciation of the audio uploaded by the terminal device according to the posterior probability distribution of the monophones corresponding to the acoustic features in each time interval in the target acoustic feature sequence.
[0149] The inventors of the present application conducted experiments to compare the model recognition effect of the HMM-DNN acoustic model using triphones as modeling units in the related art with the model recognition effect of the Chain model using diphone states as modeling units in the embodiments of the present application, and obtained the model recognition effect comparison results shown in Table 1. In order to ensure the fairness of the comparison, both the HMM-DNN model and the Chain model were trained using 380 hours of Chinese primary school students' spoken English recordings, and were tested using 10 hours of speech data.
[0150] Table 1
[0151] Chain Model HMM-DNN Model Word Error Rate (WER) 11.22 13.51
[0152] It can be found from Table 1 that the use of the Chain model can achieve a higher recognition accuracy, and theoretically, the decoding graph of the Chain model is smaller than the decoding graph of the HMM-DNN model with triphones as modeling units, and the decoding time of the Chain model is shorter.
[0153] In addition, the inventors also use the likelihood probability based on the Chain model output provided in the embodiment of the present application to perform the spoken pronunciation evaluation method and the spoken pronunciation evaluation method based on the HMM-DNN model in the related art, and score the phoneme pronunciation to evaluate the evaluation accuracy of the two methods. In order to ensure the fairness of the comparison, after completing the calculation of the phoneme posterior probability, a three-layer neural network with the same structure is used to predict the phoneme pronunciation; the input of the neural network is the phoneme posterior probability calculated by these two methods, and the output of the neural network is 1 or 0 (respectively indicating whether the phoneme pronunciation of the current evaluation is good or not). The neural network is trained using about 3000 sentences of Chinese primary school students' English oral phoneme annotation data, and uses about 1000 sentences of annotation data for testing. The indicators of experimental evaluation include recall rate (Recall), accuracy (Precision) and F value (F-measure), and the comparison results are shown in Table 2.
[0154] Table 2
[0155] Accuracy Recall F-number Related technologies 0.49 0.54 0.51 This application 0.46 0.61 0.53
[0156] It can be found from Table 2 that in terms of recall rate and F value, the evaluation effect of the method provided in the embodiment of the present application is better than the evaluation effect based on the HMM-DNN model in the related art.
[0157] The advantage of the method provided in the embodiment of the present application is that it uses the Chain model widely used in the speech recognition industry as a basis, determines the phoneme posterior probability based on the likelihood probability output by the Chain model, and performs pronunciation evaluation based on the phoneme posterior probability. On the one hand, the pronunciation evaluation effect obtained based on the Chain model is better than the pronunciation evaluation effect based on the HMM-DNN model in the related art. On the other hand, considering that the current pronunciation scoring software generally needs to maintain two models, one is the HMM-DNN model for pronunciation evaluation, and the other is the Chain model for speech recognition, which will increase the maintenance cost of the system and consume a lot of manpower and material resources; the embodiment of the present application proposes a spoken pronunciation evaluation method based on the Chain model, so that the pronunciation scoring software only needs to maintain one Chain model, which greatly reduces the maintenance cost of the software product.
[0158] With respect to the oral pronunciation evaluation method described above, the present application also provides a corresponding oral pronunciation evaluation device to enable the above-mentioned oral pronunciation evaluation method to be applied and implemented in practice.
[0159] See also Figure 7 , Figure 7 It is the above Figure 2 The structure diagram of the spoken pronunciation evaluation device 700 corresponding to the spoken pronunciation evaluation method shown in FIG. Figure 7As shown, the spoken pronunciation evaluation device 700 includes:
[0160] The audio acquisition module 701 is used to acquire the target audio to be evaluated; the target audio corresponds to the target text;
[0161] The acoustic feature extraction module 702 is used to perform acoustic feature extraction processing on the target audio to obtain a target acoustic feature sequence;
[0162] A likelihood probability determination module 703 is used to determine an acoustic likelihood probability vector according to the target acoustic feature sequence through an acoustic feature recognition model; the acoustic feature recognition model is a model that uses a diphone state as a modeling unit;
[0163] A posterior probability determination module 704, configured to determine the posterior probability of a target phoneme in the target text based on the acoustic likelihood probability vector and the target text;
[0164] The pronunciation evaluation module 705 is used to determine the target pronunciation evaluation result according to the posterior probability of the target phoneme.
[0165] Optional, in Figure 7 Based on the oral pronunciation evaluation device shown, see Figure 8 , Figure 8 FIG. 8 is a schematic diagram of another oral pronunciation evaluation device 800 provided in an embodiment of the present application. Figure 8 As shown, the posterior probability determination module 704 includes:
[0166] A forced alignment submodule 801 is used to determine, based on the acoustic likelihood probability vector and the target text, a time interval to which the acoustic features in the target acoustic feature sequence corresponding to the target diphone in the target text belong as a target time interval;
[0167] The posterior probability determination submodule 802 is used to determine the posterior probability of the target phoneme according to the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector.
[0168] Optional, in Figure 8 Based on the oral pronunciation evaluation device shown, the forced alignment submodule 801 is specifically used for:
[0169] Constructing a candidate diphone state sequence corresponding to the target text according to the duration of the target audio and the diphone states corresponding to each diphone in the target text;
[0170] For each of the candidate diphone state sequences, determining a reference likelihood probability corresponding to the candidate diphone state sequence based on the acoustic likelihood probability vector;
[0171] Selecting a target diphone state sequence from each of the candidate diphone state sequences according to the reference likelihood probabilities corresponding to each of the candidate diphone state sequences;
[0172] According to the target diphone state sequence, the time intervals to which the acoustic features in the target acoustic feature sequence corresponding to the diphones in the target text respectively belong are determined.
[0173] Optional, in Figure 8 Based on the oral pronunciation evaluation device shown, see Fig. 9 , Fig. 9 FIG. 9 is a schematic diagram of another oral pronunciation evaluation device 900 provided in an embodiment of the present application. Fig. 9 As shown, the posterior probability determination submodule 802 includes:
[0174] An HMM topology determination unit 901 is used to determine a reference Hidden Markov Model HMM topology according to the length of the target time interval;
[0175] A candidate diphone construction unit 902 is used to combine the monophones in pairs to obtain a plurality of candidate diphones corresponding to the acoustic features within the target time interval;
[0176] A diphone state sequence construction unit 903, configured to determine a diphone state sequence corresponding to each candidate diphone according to the diphone state corresponding to each candidate diphone and the reference HMM topology;
[0177] The posterior probability determination unit 904 is used to determine the posterior probability of the target phoneme based on each of the two-phoneme state sequences and according to the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector.
[0178] Optional, in Fig. 9 Based on the oral pronunciation evaluation device shown in FIG. 1 , the posterior probability determination unit 904 is specifically used for:
[0179] For each of the diphone state sequences, determining a reference likelihood probability of the candidate diphone corresponding to the diphone state sequence according to the likelihood probability corresponding to the acoustic features in the target time interval under the diphone state included in the diphone state sequence in the acoustic likelihood probability vector;
[0180] Determine the sum of the reference likelihood probabilities of the candidate diphones as the total reference likelihood probability;
[0181] For each of the candidate diphones, determining the posterior probability of the candidate diphone according to the reference likelihood probability of the candidate diphone and the total reference likelihood probability;
[0182] Based on the posterior probability of each of the candidate diphones, a posterior probability distribution of the diphones corresponding to the acoustic features in the target time interval is constructed as the posterior probability of the target phoneme.
[0183] Optional, in Fig. 9 Based on the oral pronunciation evaluation device shown in FIG. 1 , the posterior probability determination unit 904 is specifically used for:
[0184] For each of the diphone state sequences, determining a reference likelihood probability of the candidate diphone corresponding to the diphone state sequence according to the likelihood probability corresponding to the acoustic features in the target time interval under the diphone state included in the diphone state sequence in the acoustic likelihood probability vector;
[0185] Determine the sum of the reference likelihood probabilities of the candidate diphones as the total reference likelihood probability;
[0186] For the target diphone, the posterior probability of the target diphone is determined as the posterior probability of the target phoneme according to the reference likelihood probability of the target diphone and the total reference likelihood probability.
[0187] Optional, in Fig. 9 Based on the oral pronunciation evaluation device shown in FIG. 1 , the posterior probability determination unit 904 is specifically used for:
[0188] The posterior probability of the target diphone is determined according to the reference likelihood probability of the target diphone, the prior probability of the target diphone, the total reference likelihood probability, and the prior probability of each of the candidate diphones.
[0189] Optional, in Fig. 9 Based on the oral pronunciation evaluation device shown in the figure, the front monophone and the back monophone of the candidate two phonemes are used as the front phoneme and the back phoneme respectively; the posterior probability determination unit 904 is specifically used for:
[0190] For the diphone state sequence corresponding to each of the candidate diphones including the same postphone, determine the reference likelihood probability of the candidate diphone according to the likelihood probability corresponding to the acoustic features in the target time interval under the diphone state included in the diphone state sequence in the acoustic likelihood probability vector; and select the maximum reference likelihood probability from the reference likelihood probabilities of the candidate diphones including the postphone as the reference likelihood probability of the postphone;
[0191] Determine the sum of the reference likelihood probabilities of the post-phonemes as the total reference likelihood probability;
[0192] For each of the post-phonemes, determining a posterior probability of the post-phoneme according to the reference likelihood probability of the post-phoneme and the total reference likelihood probability;
[0193] Based on the posterior probabilities of the respective posterior phonemes, a posterior probability distribution of the single phoneme corresponding to the acoustic features in the target time interval is constructed as the posterior probability of the target phoneme.
[0194] Optional, in Fig. 9 Based on the oral pronunciation evaluation device shown in the figure, the front monophone and the back monophone of the candidate two phonemes are used as the front phoneme and the back phoneme respectively; the posterior probability determination unit 904 is specifically used for:
[0195] For the diphone state sequence corresponding to each of the candidate diphones including the same postphone, determine the reference likelihood probability of the candidate diphone according to the likelihood probability corresponding to the acoustic features in the target time interval under the diphone state included in the diphone state sequence in the acoustic likelihood probability vector; and select the maximum reference likelihood probability from the reference likelihood probabilities of the candidate diphones including the postphone as the reference likelihood probability of the postphone;
[0196] Determine the sum of the reference likelihood probabilities of the post-phonemes as the total reference likelihood probability;
[0197] For the target post-phoneme in the target two phonemes, the posterior probability of the target post-phoneme is determined according to the reference likelihood probability of the target post-phoneme and the total reference likelihood probability as the posterior probability of the target phoneme.
[0198] Optional, in Figure 7 Based on the oral pronunciation evaluation device shown, the pronunciation evaluation module 705 is specifically used to perform at least one of the following operations:
[0199] Determining a phoneme pronunciation evaluation result according to the posterior probability of the target phoneme through a phoneme evaluation model;
[0200] Determine the word pronunciation evaluation result through the word evaluation model according to the first posterior probability set; the first posterior probability set includes: the posterior probability of each target phoneme included in the word to be evaluated in the target text;
[0201] The sentence pronunciation evaluation result is determined by the sentence evaluation model according to a second posterior probability set; the second posterior probability set includes: the posterior probability of each target phoneme included in the sentence to be evaluated in the target text.
[0202] The oral pronunciation evaluation device provided in the embodiment of the present application takes into account that the acoustic feature recognition model with a two-phone state as a modeling unit has better acoustic modeling ability and speech recognition ability than the HMM-DNN model with a three-phoneme as a modeling unit. Therefore, the acoustic feature recognition model is introduced into the oral pronunciation evaluation process; and in order to make the acoustic likelihood probability vector output by the acoustic feature recognition model suitable for oral pronunciation evaluation, the embodiment of the present application also proposes an implementation method for determining the acoustic posterior probability based on the acoustic likelihood probability. In this way, the acoustic feature recognition model with a two-phone state as a modeling unit is used for oral pronunciation evaluation, and the use of the acoustic feature recognition model for oral pronunciation evaluation can ensure that the determined pronunciation evaluation result has a high degree of accuracy, thereby effectively improving the oral pronunciation evaluation effect.
[0203] The embodiment of the present application also provides a device for evaluating spoken pronunciation, which may specifically be a terminal device or a server. The terminal device and server provided in the embodiment of the present application will be introduced below from the perspective of hardware instantiation.
[0204] See also Fig.10 , Fig.10 Schematic diagram of the structure of the terminal device provided in the embodiment of the present application. Fig.10 For the sake of convenience, only the parts related to the embodiments of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The terminal can be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (English full name: Personal Digital Assistant, English abbreviation: PDA), a sales terminal (English full name: Point of Sales, English abbreviation: POS), a car computer, etc., taking the terminal as a smart phone as an example:
[0205] Fig.10 FIG. 1 is a block diagram showing a partial structure of a smart phone related to a terminal provided in an embodiment of the present application. Fig.10 The smart phone includes: a radio frequency (RF) circuit 1010, a memory 1020, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a wireless fidelity (WiFi) module 1070, a processor 1080, and a power supply 1090. Those skilled in the art will appreciate that Fig.10 The structure of the smartphone shown in the figure does not constitute a limitation of the smartphone, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0206] The memory 1020 can be used to store software programs and modules. The processor 1080 executes various functional applications and data processing of the smartphone by running the software programs and modules stored in the memory 1020. The memory 1020 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the smartphone (such as audio data, a phone book, etc.), etc. In addition, the memory 1020 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0207] The processor 1080 is the control center of the smartphone, which uses various interfaces and lines to connect various parts of the entire smartphone, and executes various functions of the smartphone and processes data by running or executing software programs and / or modules stored in the memory 1020, and calling data stored in the memory 1020. Optionally, the processor 1080 may include one or more processing units; preferably, the processor 1080 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 1080.
[0208] In the embodiment of the present application, the processor 1080 included in the terminal also has the following functions:
[0209] Acquire a target audio to be evaluated; the target audio corresponds to a target text;
[0210] Performing acoustic feature extraction processing on the target audio to obtain a target acoustic feature sequence;
[0211] Determining an acoustic likelihood probability vector according to the target acoustic feature sequence through an acoustic feature recognition model; the acoustic feature recognition model is a model that uses a diphone state as a modeling unit;
[0212] Determining a posterior probability of a target phoneme in the target text based on the acoustic likelihood probability vector and the target text;
[0213] The target pronunciation evaluation result is determined according to the posterior probability of the target phoneme.
[0214] Optionally, the processor 1080 is further configured to execute steps of any implementation of the spoken pronunciation evaluation method provided in the embodiments of the present application.
[0215] See also Fig.11 , Fig.11 A schematic diagram of the structure of a server 1100 provided for an embodiment of the present application. The server 1100 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 1122 (for example, one or more processors) and a memory 1132, and one or more storage media 1130 (for example, one or more mass storage devices) storing application programs 1142 or data 1144. Among them, the memory 1132 and the storage medium 1130 may be temporary storage or permanent storage. The program stored in the storage medium 1130 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1122 may be configured to communicate with the storage medium 1130 to execute a series of instruction operations in the storage medium 1130 on the server 1100.
[0216] The server 1100 may also include one or more power supplies 1126, one or more wired or wireless network interfaces 1150, one or more input and output interfaces 1158, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0217] The steps performed by the server in the above embodiment can be based on the Fig.11 The server structure shown.
[0218] The CPU 1122 is used to execute the following steps:
[0219] Acquire a target audio to be evaluated; the target audio corresponds to a target text;
[0220] Performing acoustic feature extraction processing on the target audio to obtain a target acoustic feature sequence;
[0221] Determining an acoustic likelihood probability vector according to the target acoustic feature sequence through an acoustic feature recognition model; the acoustic feature recognition model is a model that uses a diphone state as a modeling unit;
[0222] Determining a posterior probability of a target phoneme in the target text based on the acoustic likelihood probability vector and the target text;
[0223] The target pronunciation evaluation result is determined according to the posterior probability of the target phoneme.
[0224] Optionally, the CPU 1122 may also be used to execute steps of any implementation of the spoken pronunciation evaluation method provided in the embodiments of the present application.
[0225] The embodiment of the present application also provides a computer-readable storage medium for storing a computer program, wherein the computer program is used to execute any one of the implementation methods of the oral pronunciation evaluation method described in the aforementioned embodiments.
[0226] The embodiment of the present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes any one of the implementations of the oral pronunciation evaluation method described in the above embodiments.
[0227] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0228] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0229] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0230] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0231] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (full name in English: Read-Only Memory, English abbreviation: ROM), random access memory (full name in English: Random Access Memory, English abbreviation: RAM), disk or optical disk and other media that can store computer programs.
[0232] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0233] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for evaluating spoken pronunciation, It is characterized in that The method comprises: Acquire a target audio to be evaluated; the target audio corresponds to a target text; Performing acoustic feature extraction processing on the target audio to obtain a target acoustic feature sequence; Determining an acoustic likelihood probability vector according to the target acoustic feature sequence through an acoustic feature recognition model; the acoustic feature recognition model is a model that uses a diphone state as a modeling unit; Based on the acoustic likelihood probability vector and the target text, determining the time interval to which the acoustic features in the target acoustic feature sequence corresponding to the target diphone in the target text belong as the target time interval; Determining the posterior probability of the target phoneme according to the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector; Determining a target pronunciation evaluation result according to the posterior probability of the target phoneme; The step of determining, based on the acoustic likelihood probability vector and the target text, the time interval to which the acoustic features in the target acoustic feature sequence corresponding to the target diphone in the target text belong includes: Constructing a candidate diphone state sequence corresponding to the target text according to the duration of the target audio and the diphone states corresponding to each diphone in the target text; For each of the candidate diphone state sequences, determining a reference likelihood probability corresponding to the candidate diphone state sequence based on the acoustic likelihood probability vector; Selecting a target diphone state sequence from each of the candidate diphone state sequences according to the reference likelihood probabilities corresponding to each of the candidate diphone state sequences; According to the target diphone state sequence, the time intervals to which the acoustic features in the target acoustic feature sequence corresponding to the diphones in the target text respectively belong are determined.
2. The method according to claim 1, It is characterized in that Determining the posterior probability of the target phoneme according to the likelihood probability corresponding to the acoustic feature in the target time interval in the acoustic likelihood probability vector includes: Determining a reference Hidden Markov Model (HMM) topology according to the length of the target time interval; Combining the monophones in pairs to obtain a plurality of candidate diphones corresponding to the acoustic features within the target time interval; Determining a diphone state sequence corresponding to each candidate diphone according to the diphone state corresponding to each candidate diphone and the reference HMM topology; Based on each of the two-phoneme state sequences, the posterior probability of the target phoneme is determined according to the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector.
3. The method according to claim 2, It is characterized in that The determining, based on each of the two-phoneme state sequences and according to the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector, the posterior probability of the target phoneme comprises: For each of the diphone state sequences, determining a reference likelihood probability of the candidate diphone corresponding to the diphone state sequence according to the likelihood probability corresponding to the acoustic features in the target time interval under the diphone state included in the diphone state sequence in the acoustic likelihood probability vector; Determine the sum of the reference likelihood probabilities of the candidate diphones as the total reference likelihood probability; For each of the candidate diphones, determining the posterior probability of the candidate diphone according to the reference likelihood probability of the candidate diphone and the total reference likelihood probability; Based on the posterior probability of each of the candidate diphones, a posterior probability distribution of the diphones corresponding to the acoustic features in the target time interval is constructed as the posterior probability of the target phoneme.
4. The method according to claim 2, It is characterized in that The determining, based on each of the two-phoneme state sequences and according to the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector, the posterior probability of the target phoneme comprises: For each of the diphone state sequences, determining a reference likelihood probability of the candidate diphone corresponding to the diphone state sequence according to the likelihood probability corresponding to the acoustic features in the target time interval under the diphone state included in the diphone state sequence in the acoustic likelihood probability vector; Determine the sum of the reference likelihood probabilities of the candidate diphones as the total reference likelihood probability; For the target diphone, the posterior probability of the target diphone is determined as the posterior probability of the target phoneme according to the reference likelihood probability of the target diphone and the total reference likelihood probability.
5. The method according to claim 4, It is characterized in that Determining the posterior probability of the target diphone according to the reference likelihood probability of the target diphone and the total reference likelihood probability includes: The posterior probability of the target diphone is determined according to the reference likelihood probability of the target diphone, the prior probability of the target diphone, the total reference likelihood probability, and the prior probability of each of the candidate diphones.
6. The method according to claim 2, It is characterized in that The method comprises: taking the front monophone and the back monophone of the candidate two phonemes as the front phoneme and the back phoneme respectively; determining the posterior probability of the target phoneme based on the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector based on each of the two phoneme state sequences, including: For the diphone state sequence corresponding to each of the candidate diphones including the same postphone, determine the reference likelihood probability of the candidate diphone according to the likelihood probability corresponding to the acoustic features in the target time interval under the diphone state included in the diphone state sequence in the acoustic likelihood probability vector; and select the maximum reference likelihood probability from the reference likelihood probabilities of the candidate diphones including the postphone as the reference likelihood probability of the postphone; Determine the sum of the reference likelihood probabilities of the post-phonemes as the total reference likelihood probability; For each of the post-phonemes, determining a posterior probability of the post-phoneme according to the reference likelihood probability of the post-phoneme and the total reference likelihood probability; Based on the posterior probabilities of the respective posterior phonemes, a posterior probability distribution of the single phoneme corresponding to the acoustic features in the target time interval is constructed as the posterior probability of the target phoneme.
7. The method according to claim 2, It is characterized in that The method comprises: taking the front monophone and the back monophone of the candidate two phonemes as the front phoneme and the back phoneme respectively; determining the posterior probability of the target phoneme based on the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector based on each of the two phoneme state sequences, including: For the diphone state sequence corresponding to each of the candidate diphones including the same postphone, determine the reference likelihood probability of the candidate diphone according to the likelihood probability corresponding to the acoustic features in the target time interval under the diphone state included in the diphone state sequence in the acoustic likelihood probability vector; and select the maximum reference likelihood probability from the reference likelihood probabilities of the candidate diphones including the postphone as the reference likelihood probability of the postphone; Determine the sum of the reference likelihood probabilities of the post-phonemes as the total reference likelihood probability; For the target post-phoneme in the target two phonemes, the posterior probability of the target post-phoneme is determined according to the reference likelihood probability of the target post-phoneme and the total reference likelihood probability as the posterior probability of the target phoneme.
8. The method according to any one of claims 1 to 6, It is characterized in that Determining the target pronunciation evaluation result according to the posterior probability of the target phoneme includes at least one of the following: Determining a phoneme pronunciation evaluation result according to the posterior probability of the target phoneme through a phoneme evaluation model; Determining a word pronunciation evaluation result according to a first posterior probability set through a word evaluation model; The first posterior probability set includes: the posterior probability of each target phoneme included in the to-be-evaluated word in the target text; The sentence pronunciation evaluation result is determined by the sentence evaluation model according to a second posterior probability set; the second posterior probability set includes: the posterior probability of each target phoneme included in the sentence to be evaluated in the target text.
9. A spoken pronunciation evaluation device, It is characterized in that The device comprises: An audio acquisition module, used to acquire a target audio to be evaluated; the target audio corresponds to a target text; An acoustic feature extraction module, used to perform acoustic feature extraction processing on the target audio to obtain a target acoustic feature sequence; A likelihood probability determination module, used to determine an acoustic likelihood probability vector according to the target acoustic feature sequence through an acoustic feature recognition model; the acoustic feature recognition model is a model using a diphone state as a modeling unit; A posterior probability determination module, used to determine the posterior probability of the target phoneme in the target text based on the acoustic likelihood probability vector and the target text; A pronunciation evaluation module, used to determine a target pronunciation evaluation result according to the posterior probability of the target phoneme; The posterior probability determination module comprises: A forced alignment submodule, used for determining, based on the acoustic likelihood probability vector and the target text, a time interval to which the acoustic features in the target acoustic feature sequence corresponding to the target diphone in the target text belong as a target time interval; a posterior probability determination submodule, used to determine the posterior probability of the target phoneme according to the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector; The forced alignment submodule is specifically used for: Constructing a candidate diphone state sequence corresponding to the target text according to the duration of the target audio and the diphone states corresponding to each diphone in the target text; For each of the candidate diphone state sequences, determining a reference likelihood probability corresponding to the candidate diphone state sequence based on the acoustic likelihood probability vector; Selecting a target diphone state sequence from each of the candidate diphone state sequences according to the reference likelihood probabilities corresponding to each of the candidate diphone state sequences; According to the target diphone state sequence, the time intervals to which the acoustic features in the target acoustic feature sequence corresponding to the diphones in the target text respectively belong are determined.
10. The device according to claim 9, It is characterized in that The posterior probability determination submodule includes: An HMM topology determination unit is used to determine a reference Hidden Markov Model HMM topology according to the length of the target time interval; A candidate diphone construction unit, used for combining the monophones in pairs to obtain a plurality of candidate diphones corresponding to the acoustic features within the target time interval; A diphone state sequence construction unit, configured to determine a diphone state sequence corresponding to each candidate diphone according to the diphone state corresponding to each candidate diphone and the reference HMM topology; The posterior probability determination unit is used to determine the posterior probability of the target phoneme based on each of the two-phoneme state sequences and according to the likelihood probability corresponding to the acoustic features in the target time interval in the acoustic likelihood probability vector.
11. The device according to claim 10, It is characterized in that The posterior probability determination unit is specifically used for: For each of the diphone state sequences, determining a reference likelihood probability of the candidate diphone corresponding to the diphone state sequence according to the likelihood probability corresponding to the acoustic features in the target time interval under the diphone state included in the diphone state sequence in the acoustic likelihood probability vector; Determine the sum of the reference likelihood probabilities of the candidate diphones as the total reference likelihood probability; For each of the candidate diphones, determining the posterior probability of the candidate diphone according to the reference likelihood probability of the candidate diphone and the total reference likelihood probability; Based on the posterior probability of each of the candidate diphones, a posterior probability distribution of the diphones corresponding to the acoustic features in the target time interval is constructed as the posterior probability of the target phoneme.
12. The device according to claim 10, It is characterized in that The posterior probability determination unit is specifically used for: For each of the diphone state sequences, determining a reference likelihood probability of the candidate diphone corresponding to the diphone state sequence according to the likelihood probability corresponding to the acoustic features in the target time interval under the diphone state included in the diphone state sequence in the acoustic likelihood probability vector; Determine the sum of the reference likelihood probabilities of the candidate diphones as the total reference likelihood probability; For the target diphone, the posterior probability of the target diphone is determined as the posterior probability of the target phoneme according to the reference likelihood probability of the target diphone and the total reference likelihood probability.
13. The device according to claim 12, It is characterized in that The posterior probability determination unit is specifically used for: The posterior probability of the target diphone is determined according to the reference likelihood probability of the target diphone, the prior probability of the target diphone, the total reference likelihood probability, and the prior probability of each of the candidate diphones.
14. The device according to claim 10, It is characterized in that The front monophone and the back monophone of the candidate two phonemes are used as the front phoneme and the back phoneme respectively; the posterior probability determination unit is specifically used for: For the diphone state sequence corresponding to each of the candidate diphones including the same postphone, determine the reference likelihood probability of the candidate diphone according to the likelihood probability corresponding to the acoustic features in the target time interval under the diphone state included in the diphone state sequence in the acoustic likelihood probability vector; and select the maximum reference likelihood probability from the reference likelihood probabilities of the candidate diphones including the postphone as the reference likelihood probability of the postphone; Determine the sum of the reference likelihood probabilities of the post-phonemes as the total reference likelihood probability; For each of the post-phonemes, determining a posterior probability of the post-phoneme according to the reference likelihood probability of the post-phoneme and the total reference likelihood probability; Based on the posterior probabilities of the respective posterior phonemes, a posterior probability distribution of the single phoneme corresponding to the acoustic features in the target time interval is constructed as the posterior probability of the target phoneme.
15. The device according to claim 10, It is characterized in that The front monophone and the back monophone of the candidate two phonemes are used as the front phoneme and the back phoneme respectively; the posterior probability determination unit is specifically used for: For the diphone state sequence corresponding to each of the candidate diphones including the same postphone, determine the reference likelihood probability of the candidate diphone according to the likelihood probability corresponding to the acoustic features in the target time interval under the diphone state included in the diphone state sequence in the acoustic likelihood probability vector; and select the maximum reference likelihood probability from the reference likelihood probabilities of the candidate diphones including the postphone as the reference likelihood probability of the postphone; Determine the sum of the reference likelihood probabilities of the post-phonemes as the total reference likelihood probability; For the target post-phoneme in the target two phonemes, the posterior probability of the target post-phoneme is determined according to the reference likelihood probability of the target post-phoneme and the total reference likelihood probability as the posterior probability of the target phoneme.
16. The device according to any one of claims 9 to 15, It is characterized in that The pronunciation evaluation module is specifically used to perform at least one of the following operations: Determining a phoneme pronunciation evaluation result according to the posterior probability of the target phoneme through a phoneme evaluation model; Determining a word pronunciation evaluation result according to a first posterior probability set through a word evaluation model; The first posterior probability set includes: the posterior probability of each target phoneme included in the to-be-evaluated word in the target text; The sentence pronunciation evaluation result is determined by the sentence evaluation model according to a second posterior probability set; the second posterior probability set includes: the posterior probability of each target phoneme included in the sentence to be evaluated in the target text.
17. A device, It is characterized in that The device includes a processor and a memory; The memory is used to store computer programs; The processor is used to execute the spoken pronunciation evaluation method according to any one of claims 1 to 8 according to the computer program.
18. A computer-readable storage medium, It is characterized in that The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the spoken pronunciation evaluation method according to any one of claims 1 to 8.
19. A computer program product, It is characterized in that The computer program product comprises instructions, and when the instructions are executed on a computer device, the computer device executes the spoken pronunciation evaluation method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method and device for evaluating pronunciation quality, electronic device and storage medium
CN109545243A