Pronunciation evaluation methods, devices, readable media and electronic devices
By integrating feature information from different evaluation tasks into the pronunciation evaluation model, the problem of low evaluation accuracy caused by scarce training samples is solved, and higher pronunciation evaluation accuracy is achieved.
Patent Information
- Application Number
- CN202310460583.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-25
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2043-04-25
AI Technical Summary
In existing technologies, pronunciation evaluation models require training with a large number of manually labeled training samples, which leads to problems such as scarce training samples and low evaluation accuracy.
By acquiring target speech and text, determining target phoneme features, and inputting them into a pre-generated pronunciation evaluation model, a target neural network model trained with multiple sample sets is used to fuse feature information from different evaluation tasks to generate a more accurate evaluation model.
This improved the accuracy of the pronunciation evaluation model, making the evaluation values for the evaluation task more accurate.
Smart Images

Figure CN116504269B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of speech recognition technology, and more specifically, to a pronunciation evaluation method, apparatus, readable medium, and electronic device. Background Technology
[0002] With the development of the Internet, Internet-based language learning applications have also developed rapidly. Speech evaluation is an important technology to help self-taught language learners, which can evaluate learners' pronunciation in terms of accuracy, fluency, completeness, and rhythm.
[0003] In related technologies, pronunciation is evaluated using assessment models. Different models are used to evaluate pronunciation from different dimensions; for example, accuracy models evaluate the accuracy of pronunciation, while completeness models evaluate the completeness of pronunciation. However, these models require a large number of manually labeled training samples, which are costly and make training samples a scarce resource. With a limited number of training samples, the accuracy of the trained assessment model is relatively low, resulting in low accuracy of the pronunciation evaluation values determined by that model. Therefore, improving the accuracy of pronunciation evaluation has become an urgent problem to be solved. Summary of the Invention
[0004] This section is provided to briefly introduce the concepts, which will be described in detail in the Detailed Description section later. This section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] Firstly, this disclosure provides a pronunciation evaluation method, including:
[0006] Obtain the target speech to be evaluated and the target text corresponding to the target speech;
[0007] Based on the target speech and the target text, multiple target phoneme features corresponding to the target speech are determined. The target phoneme features are used to characterize the phoneme features of the target phonemes in the target text in the target speech. Different target phonemes correspond to different target phoneme features.
[0008] Multiple target phoneme features and at least one target evaluation task are input into a pre-generated pronunciation evaluation model to obtain the evaluation value corresponding to each target evaluation task output by the pronunciation evaluation model.
[0009] The pronunciation evaluation model is obtained by training the target neural network model through multiple sample sets. The sample sets include sample phoneme information of each sample phoneme in the sample text, sample phoneme features of each sample phoneme in the sample speech, multiple sample evaluation tasks, and sample evaluation values corresponding to each sample evaluation task. The sample phoneme information includes sample phoneme identifier and sample position information. The sample position information is used to characterize the position of the sample phoneme in the sample phoneme sequence corresponding to the sample text.
[0010] Secondly, this disclosure provides a pronunciation evaluation device, including:
[0011] The first acquisition module is used to acquire the target speech to be evaluated and the target text corresponding to the target speech;
[0012] The determining module is used to determine multiple target phoneme features corresponding to the target speech based on the target speech and the target text. The target phoneme features are used to characterize the phoneme features of the target phonemes in the target text in the target speech. Different target phonemes correspond to different target phoneme features.
[0013] The second acquisition module is used to input multiple target phoneme features and at least one target evaluation task into a pre-generated pronunciation evaluation model to obtain the evaluation value corresponding to each target evaluation task output by the pronunciation evaluation model.
[0014] The pronunciation evaluation model is obtained by training the target neural network model through multiple sample sets. The sample sets include sample phoneme information of each sample phoneme in the sample text, sample phoneme features of each sample phoneme in the sample speech, multiple sample evaluation tasks, and sample evaluation values corresponding to each sample evaluation task. The sample phoneme information includes sample phoneme identifier and sample position information. The sample position information is used to characterize the position of the sample phoneme in the sample phoneme sequence corresponding to the sample text.
[0015] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect of this disclosure.
[0016] Fourthly, this disclosure provides an electronic device, comprising:
[0017] A storage device on which computer programs are stored;
[0018] A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect of this disclosure.
[0019] The above technical solution obtains the target speech to be evaluated and the target text corresponding to the target speech; based on the target speech and the target text, determines multiple target phoneme features corresponding to the target speech, the target phoneme features being used to characterize the phoneme features of the target phoneme in the target text in the target speech, and different target phonemes correspond to different target phoneme features; inputs the multiple target phoneme features and at least one target evaluation task into a pre-generated pronunciation evaluation model to obtain the evaluation value corresponding to each target evaluation task output by the pronunciation evaluation model; wherein, the pronunciation evaluation model is obtained by training a target neural network model through multiple sample sets, the sample set including sample phoneme information of each sample phoneme in the sample text, sample phoneme features of each sample phoneme in the sample speech, multiple sample evaluation tasks, and sample evaluation values corresponding to each sample evaluation task, the sample phoneme information including sample phoneme identifier and sample position information, the sample position information being used to characterize the position of the sample phoneme in the sample phoneme sequence corresponding to the sample text. In other words, this disclosure simultaneously inputs multiple target phoneme features corresponding to the target speech and at least one target evaluation task into a pronunciation evaluation model, and uses the pronunciation evaluation model to evaluate the pronunciation of at least one target evaluation task. In the process of training the pronunciation evaluation model, this disclosure fuses the phoneme features corresponding to the sample speech with each evaluation task, which enables the sharing of feature information between different evaluation tasks, allowing each evaluation task to utilize more feature information, and increasing the accuracy of the generated pronunciation evaluation model. As a result, the accuracy of the evaluation value corresponding to the target evaluation task determined by the pronunciation evaluation model is also higher.
[0020] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0021] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0022] Figure 1 This is a flowchart illustrating a pronunciation evaluation method according to an exemplary embodiment of the present disclosure;
[0023] Figure 2 This is a flowchart illustrating a model training method according to an exemplary embodiment of the present disclosure;
[0024] Figure 3 This is a schematic diagram illustrating a model training method according to an exemplary embodiment of the present disclosure;
[0025] Figure 4 This is a flowchart illustrating a model training step according to an exemplary embodiment of the present disclosure;
[0026] Figure 5 This is a block diagram illustrating a pronunciation evaluation device according to an exemplary embodiment of the present disclosure;
[0027] Figure 6 This is a block diagram illustrating another pronunciation evaluation device according to an exemplary embodiment of the present disclosure;
[0028] Figure 7 This is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0029] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0030] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0031] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0032] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0033] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0034] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0035] All actions involving the acquisition of signals, information, or data in this disclosure are carried out in accordance with the relevant data protection laws and policies of the country where the location is situated, and with the authorization granted by the owner of the relevant device.
[0036] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0037] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0038] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0039] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0040] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0041] The specific embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0042] Figure 1 This is a flowchart illustrating a pronunciation evaluation method according to an exemplary embodiment of the present disclosure, such as... Figure 1 As shown, the method may include:
[0043] S101. Obtain the target speech to be evaluated and the target text corresponding to the target speech.
[0044] In this step, the target text can be displayed via an electronic device, and the user's voice signal input to the target text can be acquired via the microphone of the electronic device. The acquired voice signal is then used as the target speech. After acquiring the voice signal, it can also be preprocessed, for example, by noise reduction, to obtain the target speech. The preprocessing method can be a signal processing method in the prior art, and this disclosure does not limit it.
[0045] S102. Based on the target speech and the target text, determine multiple target phoneme features corresponding to the target speech.
[0046] Specifically, the target phoneme feature can be used to characterize the phoneme features of the target phoneme in the target text within the target speech. Different target phonemes correspond to different target phoneme features. For example, the target phoneme feature can be a GoP (Goodness of Pronunciation) feature.
[0047] In this step, after obtaining the target speech and its corresponding target text, the target phoneme features of each target phoneme in the target text within the target speech can be determined based on the target speech and the target text. Taking GoP features as an example, if the target text is "Its Name", then the target phoneme sequence corresponding to the target text is "IH TS|N EY M". This target phoneme sequence contains 6 target phonemes, and the multiple target phoneme features corresponding to the target speech can be represented as "GoP[IH], GoP[T], GoP[S], GoP[N], GoP[EY], GoP[M]".
[0048] In one possible implementation, the target speech and the target text can be input into a pre-generated phoneme feature acquisition model to obtain multiple target phoneme features output by the phoneme feature acquisition model. The phoneme feature acquisition model can be an existing ASR (Automatic Speech Recognition) acoustic model.
[0049] S103. Input multiple target phoneme features and at least one target evaluation task into a pre-generated pronunciation evaluation model to obtain the evaluation value corresponding to each target evaluation task output by the pronunciation evaluation model.
[0050] The pronunciation evaluation model can be obtained by training the target neural network model through multiple sample sets. The sample set includes sample phoneme information of each sample phoneme in the sample text, sample phoneme features of each sample phoneme in the sample speech, multiple sample evaluation tasks, and sample evaluation values corresponding to each sample evaluation task. The sample phoneme information includes sample phoneme identifier and sample position information. The sample position information is used to characterize the position of the sample phoneme in the sample phoneme sequence corresponding to the sample text.
[0051] In this step, after determining multiple target phoneme features corresponding to the target speech, at least one target evaluation task that needs to be evaluated for the target speech can be determined. In one possible implementation, the target evaluation task can be multiple preset evaluation tasks that the pronunciation evaluation model can perform. These preset evaluation tasks can be evaluation tasks that the pronunciation evaluation model can perform for pronunciation evaluation. For example, the preset evaluation tasks can include phoneme evaluation, word evaluation, and sentence evaluation. For phoneme evaluation and word evaluation, accuracy evaluation can be included. For sentence evaluation, fluency evaluation, completeness evaluation, prosody evaluation, etc., are not limited in this disclosure.
[0052] In another possible implementation, at least one target evaluation task can be determined from multiple preset evaluation tasks based on the number of phonemes in the target text selected by the user. For example, if the user selects a target phoneme in the target text, the target evaluation task can be phoneme evaluation, that is, evaluating the accuracy of the pronunciation of the target phoneme in the target speech. Continuing with the target text in step S102 as an example, if the user selects the target phoneme "S", the target evaluation task is to evaluate the accuracy of the target phoneme "S"; if the user selects the target word "Name", the target evaluation task is to evaluate the accuracy of the target word "Name"; if the user selects all the target phonemes in the target text, the target evaluation task can be sentence evaluation, that is, evaluating the fluency, completeness, and prosody of the entire sentence of the target speech.
[0053] After determining at least one target evaluation task, multiple target phoneme features and at least one target evaluation task can be input into the pronunciation evaluation model. For each target evaluation task, the pronunciation evaluation model can extract features from multiple target phoneme features to obtain the target phoneme evaluation features corresponding to the target evaluation task. The target phoneme evaluation features are then pooled according to the target evaluation task to obtain the target pooled features corresponding to the target evaluation task. The evaluation value corresponding to the target evaluation task is then determined based on the target pooled features.
[0054] Using the above method, multiple target phoneme features corresponding to the target speech and at least one target evaluation task are simultaneously input into the pronunciation evaluation model. The pronunciation evaluation model is then used to evaluate the pronunciation of at least one target evaluation task. In the process of training the pronunciation evaluation model, this disclosure integrates the phoneme features corresponding to the sample speech with each evaluation task, which enables the sharing of feature information between different evaluation tasks. This allows each evaluation task to utilize more feature information, resulting in a higher accuracy of the generated pronunciation evaluation model. Consequently, the accuracy of the evaluation value corresponding to the target evaluation task determined by the pronunciation evaluation model is also higher.
[0055] Figure 2 This is a flowchart illustrating a model training method according to an exemplary embodiment of the present disclosure, such as... Figure 2 As shown, the method may include:
[0056] S21. Obtain multiple sample sets and determine the current sample set from these multiple sample sets.
[0057] In this step, multiple sample texts and corresponding sample speech for each sample text can be acquired. The sample speech can be collected from multiple users regarding the sample text. For each sample text, the sample phoneme information for each sample phoneme in the sample text, the sample phoneme features of each sample phoneme in the sample speech, multiple sample evaluation tasks, and the sample evaluation value corresponding to each sample evaluation task can be determined. After obtaining multiple sample sets, any sample set from the multiple sample sets can be used as the current sample set.
[0058] For example, taking any sample set as an example, the sample text and the corresponding sample speech of the sample text can be input into the phoneme feature acquisition model to obtain multiple sample phoneme features output by the phoneme feature acquisition model. For each sample phoneme in the sample text, the sample phoneme identifier of the sample phoneme can be determined through a pre-set identifier association relationship. The identifier association relationship can include the correspondence between different phonemes and phoneme identifiers.
[0059] Figure 3 This is a schematic diagram illustrating a model training method according to an exemplary embodiment of the present disclosure, such as... Figure 3As shown, if the sample text is "Its Name", then the corresponding sample phoneme sequence is "IH TS|N EY M", which contains 6 sample phonemes. After inputting the sample text and the sample speech into the phoneme feature acquisition model, multiple sample phoneme features corresponding to the sample speech are obtained: "GoP[IH], GoP[T], GoP[S], GoP[N], GoP[EY], GoP[M]". The sample phoneme identifier of the sample phoneme "IH" can be represented as "Phn[IH]", the sample phoneme identifier of the sample phoneme "T" can be represented as "Phn[T]", the sample phoneme identifier of the sample phoneme "S" can be represented as "Phn[S]", the sample phoneme identifier of the sample phoneme "N" can be represented as "Phn[N]", the sample phoneme identifier of the sample phoneme "EY" can be represented as "Phn[EY]", and the sample phoneme identifier of the sample phoneme "M" can be represented as "Phn[M]". For each sample phoneme in the sample text, the sample position information of the sample phoneme can be determined according to the position of the sample phoneme in the sample phoneme sequence of the sample text. For example, the sample position information of the sample phoneme "IH" can be represented as "Pos[0]", the sample phoneme identifier of the sample phoneme "T" can be represented as "Pos[1]", the sample phoneme identifier of the sample phoneme "S" can be represented as "Pos[2]", the sample phoneme identifier of the sample phoneme "N" can be represented as "Pos[3]", the sample phoneme identifier of the sample phoneme "EY" can be represented as "Pos[4]", and the sample phoneme identifier of the sample phoneme "M" can be represented as "Pos[5]".
[0060] It should be noted that the evaluation tasks for multiple samples in different sample sets can be the same or different, and this disclosure does not impose any restrictions on this. For each sample set, multiple evaluation tasks for the sample speech in that sample set can be manually labeled to obtain the sample evaluation value corresponding to each evaluation task.
[0061] S22. Based on the current sample set, repeatedly execute the model training steps until the trained target neural network model satisfies the preset stopping iteration condition based on multiple sample evaluation values and multiple current prediction evaluation values of the current sample set. Then, use the trained target neural network model as the pronunciation evaluation model.
[0062] Wherein, the current predicted evaluation value is the evaluation value corresponding to the sample evaluation task output by the trained target neural network model after the current sample set is input. The preset stopping iteration condition can be any stopping iteration condition in the prior art, and this disclosure does not limit it.
[0063] In this step, the model training step can be performed based on the current sample set. The target neural network outputs the current predicted evaluation value corresponding to each sample evaluation task. Based on multiple sample evaluation values and multiple current predicted evaluation values, it is determined whether the target neural network meets the preset stopping iteration condition. If it is determined that the target neural network meets the preset stopping iteration condition, the target neural network model can be used as the pronunciation evaluation model. If it is determined that the target neural network does not meet the preset stopping iteration condition, a new current sample set can be determined from multiple sample sets, and the model training step can continue to be performed based on the new current sample set until it is determined that the trained target neural network model meets the preset stopping iteration condition.
[0064] Figure 4 This is a flowchart illustrating a model training step according to an exemplary embodiment of the present disclosure, such as... Figure 4 As shown, the model training steps may include:
[0065] S1. Obtain the current predicted evaluation value for each evaluation task in the current sample set through the target neural network model.
[0066] For example, the current sample set can be input into the target neural network model to obtain the current predicted evaluation value corresponding to each sample evaluation task output by the target neural network model.
[0067] In one possible implementation, for each sample evaluation task, the target neural network model can be used to extract features from multiple current sample phoneme features and multiple current sample phoneme information of the current sample set to obtain the phoneme evaluation features corresponding to the sample evaluation task. Based on the phoneme evaluation features and the sample evaluation task, the current predicted evaluation value corresponding to the sample evaluation task can be determined.
[0068] The target neural network model may include a feature extraction sub-model. For each current sample phoneme in the current sample set, the current sample phoneme features and current sample phoneme information of the current sample phoneme can be concatenated to obtain the current sample concatenation information. According to the sample evaluation task, the feature extraction sub-model extracts features from the current sample concatenation information to obtain the phoneme evaluation features corresponding to the sample evaluation task.
[0069] For example, continuing with Figure 3Taking the model training diagram shown as an example, the current sample phoneme sequence corresponding to the current sample text is "IH TS|N EY M". For the current sample phoneme "IH" in the current sample phoneme sequence, the current sample phoneme feature "GoP[IH]", the current sample phoneme identifier "Phn[T]", and the current sample position information "Pos[0]" of the current sample phoneme "IH" can be concatenated to obtain the current sample concatenation information of the current sample phoneme "IH". Referring to the above processing method for the current sample phoneme "IH" in the current sample phoneme sequence, other current sample phonemes in the current sample phoneme sequence can be processed to obtain the current sample concatenation information corresponding to each current sample phoneme in the current sample phoneme sequence, that is, to obtain 6 current sample concatenation information.
[0070] After determining the current sample concatenation information corresponding to each current sample phoneme in the current sample set, for each sample evaluation task in the current sample set, the sample evaluation task can be concatenated with multiple current sample concatenation information to obtain the current sample evaluation information. The current sample evaluation information is then input into the feature extraction sub-model, and feature extraction is performed through the feature extraction sub-model to obtain the phoneme evaluation features corresponding to the sample evaluation task.
[0071] In one possible implementation, the target neural network model includes a pooling layer and an evaluation sub-model. The output of the pooling layer is coupled to the input of the evaluation sub-model. For each sample evaluation task, after determining the phoneme evaluation features corresponding to each sample evaluation task, a target sample phoneme can be determined from multiple current sample phonemes in the current sample set according to the sample evaluation task. Based on the target sample phoneme, the phoneme evaluation features are pooled to obtain the sample pooling features corresponding to the sample evaluation task. Based on the sample pooling features, the current predicted evaluation value corresponding to the sample evaluation task is determined through the evaluation sub-model.
[0072] Specifically, the range of sample phonemes corresponding to the sample evaluation task can be determined by pre-set range association relationships, which include the correspondence between different evaluation tasks and phoneme ranges; based on the sample phoneme range, the target sample phoneme can be determined from multiple current sample phonemes.
[0073] Continuing with the example of the current sample text "Its Name", if the evaluation task is phoneme evaluation, the range of phonemes in the sample can be [0,0], [1,1], [2,2], [3,3], [4,4], [5,5], representing the evaluation of the accuracy of each phoneme in the current sample speech; if the evaluation task is word evaluation, the range of phonemes in the sample can be [0,2], [3,5], representing the evaluation of the accuracy of each word in the current sample speech; if the evaluation task is sentence evaluation, the range of phonemes in the sample can be [0,5], representing the evaluation of the fluency, completeness, and prosody of the entire sentence in the current sample speech.
[0074] After determining the range of sample phonemes corresponding to each sample evaluation task, the boundary range corresponding to that sample evaluation task can be determined based on the range of sample phonemes. Then, the phoneme evaluation features corresponding to that sample evaluation task are pooled based on this boundary range. For example, pooling can be performed using existing techniques such as boundary pooling or average pooling to obtain the sample pooling features corresponding to that sample evaluation task. Afterwards, the evaluation sub-model can be used to score based on these pooling features to obtain the current predicted evaluation value for that sample evaluation task. For example, this evaluation sub-model can be a linear mapping layer.
[0075] S2. If the target neural network model does not meet the preset stopping iteration condition based on multiple current prediction evaluation values and multiple sample evaluation values, determine the target loss value based on multiple current prediction evaluation values and multiple sample evaluation values, update the parameters of the target neural network model based on the target loss value, obtain the trained target neural network model, use the trained target neural network model as the new target neural network model, and determine a new current sample set from multiple sample sets.
[0076] For example, after determining multiple current predicted evaluation values, a target loss value can be determined based on these multiple current predicted evaluation values and multiple sample evaluation values. For instance, for each sample evaluation task, the task loss value corresponding to that sample evaluation task can be determined based on the current predicted evaluation value and the sample evaluation value corresponding to that sample evaluation task. The target loss value can be the average of multiple task loss values, the maximum task loss value among the multiple task loss values, or different weights can be set for different sample evaluation tasks, and the target loss value can be determined based on these weights and the task loss value. The above methods for determining the target loss value are illustrative examples, and this disclosure does not limit them.
[0077] After determining the target loss value, it can be determined whether the target loss value is less than or equal to a preset loss value threshold. If the target loss value is less than or equal to the preset loss value threshold, it can be determined that the target neural network model meets the preset stopping iteration condition. If the target loss value is greater than the preset loss value threshold, it can be determined that the target neural network model does not meet the preset loss value threshold. The parameters of the target neural network model are updated according to the target loss value to obtain the trained target neural network model. The trained target neural network model is used as the new target neural network model, and a new current sample set is determined from multiple sample sets. Based on the new current sample set, the above steps S1 to S2 are continued.
[0078] Using the above-mentioned model training method, during the training of the pronunciation evaluation model, the feature extraction sub-model not only extracts the feature information of phonemes, but also extracts the feature information of the evaluation task. In this way, the feature extraction sub-model can utilize the feature information between different evaluation tasks, so that the feature information between different evaluation tasks is complementary, thereby improving the accuracy of the pronunciation evaluation model. Furthermore, since the pronunciation evaluation model can make full use of the feature information between different evaluation tasks, it reduces the dependence on the number of training samples, and can generate a pronunciation evaluation model with relatively high accuracy even when the number of training samples is limited.
[0079] Figure 5 This is a block diagram illustrating a pronunciation evaluation device according to an exemplary embodiment of the present disclosure, such as... Figure 5 As shown, the device may include:
[0080] The first acquisition module 501 is used to acquire the target speech to be evaluated and the target text corresponding to the target speech;
[0081] The determining module 502 is used to determine multiple target phoneme features corresponding to the target speech based on the target speech and the target text. The target phoneme features are used to characterize the phoneme features of the target phonemes in the target text in the target speech. Different target phonemes correspond to different target phoneme features.
[0082] The second acquisition module 503 is used to input multiple target phoneme features and at least one target evaluation task into a pre-generated pronunciation evaluation model to obtain the evaluation value corresponding to each target evaluation task output by the pronunciation evaluation model.
[0083] The pronunciation evaluation model is obtained by training the target neural network model through multiple sample sets. The sample sets include sample phoneme information of each sample phoneme in the sample text, sample phoneme features of each sample phoneme in the sample speech, multiple sample evaluation tasks, and sample evaluation values corresponding to each sample evaluation task. The sample phoneme information includes sample phoneme identifier and sample position information. The sample position information is used to characterize the position of the sample phoneme in the sample phoneme sequence corresponding to the sample text.
[0084] Optionally, Figure 6 This is a block diagram illustrating another pronunciation evaluation device according to an exemplary embodiment of the present disclosure, such as... Figure 6 As shown, the device also includes:
[0085] Model training module 504 is used for:
[0086] Obtain multiple sample sets and determine the current sample set from these multiple sample sets;
[0087] Based on the current sample set, the model training steps are executed iteratively until the trained target neural network model meets the preset stopping iteration condition based on multiple sample evaluation values and multiple current predicted evaluation values of the current sample set. The trained target neural network model is then used as the pronunciation evaluation model, and the current predicted evaluation value is the evaluation value corresponding to the sample evaluation task output after the current sample set is input into the trained target neural network model.
[0088] The training steps for this model include:
[0089] The target neural network model is used to obtain the current predicted evaluation value for each evaluation task in the current sample set.
[0090] If, based on multiple current prediction evaluation values and multiple sample evaluation values, it is determined that the target neural network model does not meet the preset stopping iteration condition, a target loss value is determined based on multiple current prediction evaluation values and multiple sample evaluation values. The parameters of the target neural network model are updated based on the target loss value to obtain the trained target neural network model. The trained target neural network model is then used as the new target neural network model, and a new current sample set is determined from multiple sample sets.
[0091] Optionally, the model training module 504 is also used for:
[0092] For each sample evaluation task, based on the sample evaluation task, the target neural network model extracts features from multiple current sample phoneme features and multiple current sample phoneme information of the current sample set to obtain the phoneme evaluation features corresponding to the sample evaluation task. Based on the phoneme evaluation features and the sample evaluation task, the current predicted evaluation value corresponding to the sample evaluation task is determined.
[0093] Optionally, the target neural network model includes a feature extraction sub-model, the model training module 504, which is also used for:
[0094] For each current sample phoneme in the current sample set, the current sample phoneme features and the current sample phoneme information of the current sample phoneme are concatenated to obtain the current sample concatenation information.
[0095] Based on the sample evaluation task, the feature extraction sub-model is used to extract features from the current sample splicing information to obtain the phoneme evaluation features corresponding to the sample evaluation task.
[0096] Optionally, the target neural network model includes a pooling layer and an evaluation sub-model, the output of the pooling layer being coupled to the input of the evaluation sub-model, and the model training module 504 is further used for:
[0097] Based on the sample evaluation task, the target sample phoneme is determined from multiple current sample phonemes in the current sample set;
[0098] Based on the target sample phoneme, the evaluation features of the phoneme are pooled to obtain the sample pooling features corresponding to the sample evaluation task.
[0099] Based on the sample pooling characteristics, the current predicted evaluation value corresponding to the sample evaluation task is determined through the evaluation sub-model.
[0100] Optionally, the model training module 504 is also used for:
[0101] The range of sample phonemes corresponding to the sample evaluation task is determined by pre-set range association relationships, which include the correspondence between different evaluation tasks and phoneme ranges;
[0102] Based on the range of the sample phonemes, the target sample phoneme is determined from multiple current sample phonemes.
[0103] Optionally, the determining module 502 is also used for:
[0104] The target speech and the target text are input into a pre-generated phoneme feature acquisition model to obtain multiple target phoneme features output by the phoneme feature acquisition model.
[0105] Using the aforementioned apparatus, multiple target phoneme features corresponding to the target speech and at least one target evaluation task are simultaneously input into the pronunciation evaluation model. The pronunciation evaluation model is then used to evaluate the pronunciation of at least one target evaluation task. In the process of training the pronunciation evaluation model, this disclosure integrates the phoneme features corresponding to the sample speech with each evaluation task, enabling the sharing of feature information between different evaluation tasks. This allows each evaluation task to utilize more feature information, resulting in a higher accuracy of the generated pronunciation evaluation model. Consequently, the accuracy of the evaluation value corresponding to the target evaluation task determined based on the pronunciation evaluation model is also higher.
[0106] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0107] The following is for reference. Figure 7 The diagram illustrates a structural schematic of an electronic device 700 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0108] like Figure 7 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0109] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0110] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.
[0111] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0112] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0113] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0114] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: acquire a target speech to be evaluated and a target text corresponding to the target speech; determine multiple target phoneme features corresponding to the target speech based on the target speech and the target text, wherein the target phoneme features are used to characterize the phoneme features of target phonemes in the target text in the target speech, and different target phonemes correspond to different target phoneme features; input the multiple target phoneme features and at least one target evaluation task into a pre-generated pronunciation evaluation model to obtain the evaluation value corresponding to each target evaluation task output by the pronunciation evaluation model; wherein the pronunciation evaluation model is obtained by training a target neural network model through multiple sample sets, the sample sets including sample phoneme information of each sample phoneme in the sample text, sample phoneme features of each sample phoneme in the sample speech, multiple sample evaluation tasks, and sample evaluation values corresponding to each sample evaluation task, the sample phoneme information including sample phoneme identifier and sample position information, the sample position information being used to characterize the position of the sample phoneme in the sample phoneme sequence corresponding to the sample text.
[0115] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0116] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0117] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the modules do not necessarily limit the module itself; for example, the first acquisition module can also be described as "a module for acquiring the target speech to be evaluated and the target text corresponding to the target speech".
[0118] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0119] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0120] According to one or more embodiments of this disclosure, Example 1 provides a pronunciation evaluation method, comprising: acquiring a target speech to be evaluated and a target text corresponding to the target speech; determining multiple target phoneme features corresponding to the target speech based on the target speech and the target text, wherein the target phoneme features are used to characterize the phoneme features of target phonemes in the target text in the target speech, and different target phonemes correspond to different target phoneme features; inputting the multiple target phoneme features and at least one target evaluation task into a pre-generated pronunciation evaluation model to obtain an evaluation value corresponding to each target evaluation task output by the pronunciation evaluation model; wherein the pronunciation evaluation model is obtained by training a target neural network model through multiple sample sets, the sample sets including sample phoneme information of each sample phoneme in the sample text, sample phoneme features of each sample phoneme in the sample speech, multiple sample evaluation tasks, and sample evaluation values corresponding to each sample evaluation task, wherein the sample phoneme information includes sample phoneme identifiers and sample position information, and the sample position information is used to characterize the position of the sample phoneme in the sample phoneme sequence corresponding to the sample text.
[0121] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein the pronunciation evaluation model is pre-generated by: acquiring a plurality of sample sets and determining a current sample set from the plurality of sample sets; performing a model training step iteratively based on the current sample set until the trained target neural network model satisfies a preset stopping iteration condition based on the plurality of sample evaluation values and the plurality of current predicted evaluation values of the current sample set; using the trained target neural network model as the pronunciation evaluation model, wherein the current predicted evaluation value is the evaluation value corresponding to the sample evaluation task output by the trained target neural network model after the current sample set is input; the model training step includes: acquiring the current predicted evaluation value corresponding to each sample evaluation task in the current sample set through the target neural network model; if the target neural network model does not satisfy the preset stopping iteration condition based on the plurality of current predicted evaluation values and the plurality of sample evaluation values, determining a target loss value based on the plurality of current predicted evaluation values and the plurality of sample evaluation values; updating the parameters of the target neural network model based on the target loss value to obtain a trained target neural network model; using the trained target neural network model as a new target neural network model; and determining a new current sample set from the plurality of sample sets.
[0122] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, wherein obtaining the current predicted evaluation value corresponding to each sample evaluation task in the current sample set through the target neural network model includes: for each sample evaluation task, based on the sample evaluation task, performing feature extraction on multiple current sample phoneme features and multiple current sample phoneme information of the current sample set through the target neural network model to obtain phoneme evaluation features corresponding to the sample evaluation task, and determining the current predicted evaluation value corresponding to the sample evaluation task based on the phoneme evaluation features and the sample evaluation task.
[0123] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 3, wherein the target neural network model includes a feature extraction sub-model. The step of extracting features from multiple current sample phoneme features and multiple current sample phoneme information of the current sample set using the target neural network model to obtain phoneme evaluation features corresponding to the sample evaluation task includes: for each current sample phoneme in the current sample set, concatenating the current sample phoneme features and the current sample phoneme information of the current sample phoneme to obtain current sample concatenation information; and extracting features from the current sample concatenation information using the feature extraction sub-model to obtain phoneme evaluation features corresponding to the sample evaluation task.
[0124] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 3, wherein the target neural network model includes a pooling layer and an evaluation sub-model, the output of the pooling layer is coupled to the input of the evaluation sub-model, and the step of determining the current predicted evaluation value corresponding to the sample evaluation task based on the phoneme evaluation features and the sample evaluation task includes: determining a target sample phoneme from a plurality of current sample phonemes in the current sample set according to the sample evaluation task; performing pooling processing on the phoneme evaluation features according to the target sample phoneme to obtain sample pooling features corresponding to the sample evaluation task; and determining the current predicted evaluation value corresponding to the sample evaluation task through the evaluation sub-model based on the sample pooling features.
[0125] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 5, wherein determining the target sample phoneme from a plurality of current sample phonemes in the current sample set according to the sample evaluation task includes: determining the sample phoneme range corresponding to the sample evaluation task through a pre-set range association relationship, wherein the range association relationship includes the correspondence between different evaluation tasks and phoneme ranges; and determining the target sample phoneme from a plurality of current sample phonemes according to the sample phoneme range.
[0126] According to one or more embodiments of this disclosure, Example 7 provides a method of any of Examples 1-6, wherein determining multiple target phoneme features corresponding to the target speech based on the target speech and the target text includes: inputting the target speech and the target text into a pre-generated phoneme feature acquisition model to obtain multiple target phoneme features output by the phoneme feature acquisition model.
[0127] According to one or more embodiments of this disclosure, Example 8 provides a pronunciation evaluation device, comprising: a first acquisition module, configured to acquire a target speech to be evaluated and a target text corresponding to the target speech; a determination module, configured to determine multiple target phoneme features corresponding to the target speech based on the target speech and the target text, wherein the target phoneme features are used to characterize the phoneme features of target phonemes in the target text in the target speech, and different target phonemes correspond to different target phoneme features; a second acquisition module, configured to input the multiple target phoneme features and at least one target evaluation task into a pre-generated pronunciation evaluation model to obtain an evaluation value corresponding to each target evaluation task output by the pronunciation evaluation model; wherein the pronunciation evaluation model is obtained by training a target neural network model through multiple sample sets, the sample sets including sample phoneme information of each sample phoneme in the sample text, sample phoneme features of each sample phoneme in the sample speech, multiple sample evaluation tasks, and sample evaluation values corresponding to each sample evaluation task, the sample phoneme information including sample phoneme identifier and sample position information, the sample position information being used to characterize the position of the sample phoneme in the sample phoneme sequence corresponding to the sample text.
[0128] According to one or more embodiments of this disclosure, Example 9 provides the apparatus of Example 8, which further includes: a model training module, configured to acquire a plurality of sample sets and determine a current sample set from the plurality of sample sets; cyclically execute a model training step based on the current sample set until a trained target neural network model is determined to satisfy a preset stopping iteration condition based on a plurality of sample evaluation values and a plurality of current predicted evaluation values of the current sample set, and use the trained target neural network model as the pronunciation evaluation model, wherein the current predicted evaluation value is the evaluation value corresponding to the sample evaluation task output by the current sample set after inputting it into the trained target neural network model; the model training step includes: acquiring a current predicted evaluation value corresponding to each sample evaluation task in the current sample set through the target neural network model; if the target neural network model is determined not to satisfy the preset stopping iteration condition based on a plurality of current predicted evaluation values and a plurality of sample evaluation values, determining a target loss value based on a plurality of current predicted evaluation values and a plurality of sample evaluation values, updating the parameters of the target neural network model based on the target loss value to obtain a trained target neural network model, using the trained target neural network model as a new target neural network model, and determining a new current sample set from the plurality of sample sets.
[0129] According to one or more embodiments of this disclosure, Example 10 provides the apparatus of Example 9, wherein the model training module is further configured to: for each of the sample evaluation tasks, extract features from multiple current sample phoneme features and multiple current sample phoneme information of the current sample set through the target neural network model to obtain phoneme evaluation features corresponding to the sample evaluation task, and determine the current predicted evaluation value corresponding to the sample evaluation task based on the phoneme evaluation features and the sample evaluation task.
[0130] According to one or more embodiments of this disclosure, Example 11 provides the apparatus of Example 10, wherein the target neural network model includes a feature extraction sub-model, and the model training module is further configured to: for each current sample phoneme in the current sample set, concatenate the current sample phoneme features and the current sample phoneme information of the current sample phoneme to obtain current sample concatenation information; and, according to the sample evaluation task, extract features from the current sample concatenation information through the feature extraction sub-model to obtain phoneme evaluation features corresponding to the sample evaluation task.
[0131] According to one or more embodiments of this disclosure, Example 12 provides an apparatus of Example 10, wherein the target neural network model includes a pooling layer and an evaluation sub-model, the output of the pooling layer being coupled to the input of the evaluation sub-model, and the model training module is further configured to: determine a target sample phoneme from a plurality of current sample phonemes in the current sample set according to the sample evaluation task; perform pooling processing on the phoneme evaluation features according to the target sample phoneme to obtain sample pooling features corresponding to the sample evaluation task; and determine the current predicted evaluation value corresponding to the sample evaluation task through the evaluation sub-model based on the sample pooling features.
[0132] According to one or more embodiments of this disclosure, Example 13 provides the apparatus of Example 12, wherein the model training module is further configured to: determine the range of sample phonemes corresponding to the sample evaluation task through a pre-set range association relationship, wherein the range association relationship includes the correspondence between different evaluation tasks and phoneme ranges; and determine the target sample phoneme from a plurality of current sample phonemes according to the sample phoneme range.
[0133] According to one or more embodiments of the present disclosure, Example 14 provides an apparatus of any of Examples 8-13, wherein the determining module is further configured to: input the target speech and the target text into a pre-generated phoneme feature acquisition model to obtain a plurality of the target phoneme features output by the phoneme feature acquisition model.
[0134] According to one or more embodiments of the present disclosure, Example 15 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any of Examples 1-7.
[0135] According to one or more embodiments of this disclosure, Example 16 provides an electronic device including: a storage device having a computer program stored thereon; and a processing device for executing the computer program in the storage device to implement the steps of the method described in any of Examples 1-7.
[0136] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0137] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0138] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. A method of evaluating a pronunciation, characterized by, The method comprises: obtaining a target speech to be evaluated and a target text corresponding to the target speech; determining, according to the target speech and the target text, a plurality of target phoneme features corresponding to the target speech, the target phoneme features being used to represent phoneme features of target phonemes in the target text in the target speech, different target phonemes corresponding to different target phoneme features; inputting the plurality of target phoneme features and at least one target evaluation task into a pre-generated pronunciation evaluation model to obtain an evaluation value corresponding to each target evaluation task output by the pronunciation evaluation model; wherein the pronunciation evaluation model is obtained by training a target neural network model through a plurality of sample sets, the sample set comprising sample phoneme information of each sample phoneme of a sample text, sample phoneme features of each sample phoneme in a sample speech, a plurality of sample evaluation tasks, and a sample evaluation value corresponding to each sample evaluation task, the sample phoneme information comprising a sample phoneme identifier and sample position information used to represent the position of the sample phoneme in a sample phoneme sequence corresponding to the sample text.
2. The method of claim 1, wherein, The pronunciation evaluation model is pre-generated in the following manner: obtaining a plurality of sample sets and determining a current sample set from the plurality of sample sets; according to the current sample set, performing a model training step in a loop until a trained target neural network model meets a preset stop iteration condition according to a plurality of sample evaluation values and a plurality of current prediction evaluation values of the current sample set, the current prediction evaluation value being an evaluation value corresponding to the sample evaluation task output by the trained target neural network model after inputting the current sample set into the trained target neural network model, and taking the trained target neural network model as the pronunciation evaluation model; the model training step comprises: obtaining, by the target neural network model, a current prediction evaluation value corresponding to each sample evaluation task in the current sample set; in a case where the target neural network model does not meet the preset stop iteration condition according to a plurality of current prediction evaluation values and a plurality of sample evaluation values, determining a target loss value according to the plurality of current prediction evaluation values and the plurality of sample evaluation values, updating parameters of the target neural network model according to the target loss value to obtain a trained target neural network model, taking the trained target neural network model as a new target neural network model, and determining a new current sample set from the plurality of sample sets.
3. The method of claim 2, wherein, The target neural network model comprises: for each sample evaluation task, performing feature extraction on a plurality of current sample phoneme features of the current sample set and a plurality of current sample phoneme information of the current sample set by the target neural network model according to the sample evaluation task to obtain phoneme evaluation features corresponding to the sample evaluation task, and determining a current prediction evaluation value corresponding to the sample evaluation task according to the phoneme evaluation features and the sample evaluation task.
4. The method of claim 3, wherein, The target neural network model comprises a feature extraction sub-model, and the feature extraction sub-model is configured to extract features of the current sample phoneme features of the current sample set and the current sample phoneme information of the current sample set according to the sample evaluation task, to obtain phoneme evaluation features corresponding to the sample evaluation task, and the phoneme evaluation features comprise: The current sample phoneme feature of the current sample phoneme and the current sample phoneme information of the current sample phoneme are spliced to obtain current sample splicing information for each current sample phoneme of the current sample set; The feature extraction sub-model is configured to extract features of the current sample splicing information according to the sample evaluation task, to obtain phoneme evaluation features corresponding to the sample evaluation task.
5. The method of claim 3, wherein, The target neural network model comprises a pooling layer and an evaluation sub-model, and an output of the pooling layer is coupled with an input of the evaluation sub-model, and the current predicted evaluation value corresponding to the sample evaluation task is determined according to the phoneme evaluation features and the sample evaluation task, and the current predicted evaluation value comprises: A target sample phoneme is determined from the current sample phonemes of the current sample set according to the sample evaluation task; The phoneme evaluation features are pooled to obtain sample pooling features corresponding to the sample evaluation task according to the target sample phoneme; The current predicted evaluation value corresponding to the sample evaluation task is determined by the evaluation sub-model according to the sample pooling features.
6. The method of claim 5, wherein, The target sample phoneme is determined from the current sample phonemes of the current sample set according to the sample evaluation task, and the target sample phoneme comprises: A sample phoneme range corresponding to the sample evaluation task is determined by a pre-set range association relationship, and the range association relationship comprises a corresponding relationship between different evaluation tasks and phoneme ranges; The target sample phoneme is determined from the current sample phonemes according to the sample phoneme range.
7. The method according to any one of claims 1 to 6, characterized in that, The target phoneme features corresponding to the target speech are determined according to the target speech and the target text, and the target phoneme features comprise: The target speech and the target text are input into a pre-generated phoneme feature acquisition model to obtain the target phoneme features output by the phoneme feature acquisition model.
8. A voice quality evaluation apparatus characterized by comprising: Comprise: The first acquisition module is configured to acquire a target speech to be evaluated and a target text corresponding to the target speech; The determination module is configured to determine target phoneme features corresponding to the target speech according to the target speech and the target text, the target phoneme features are used to represent phoneme features of target phonemes in the target text in the target speech, different target phonemes correspond to different target phoneme features, and each target evaluation task corresponds to an evaluation value output by a pronunciation evaluation model. The second acquisition module is configured to input the target phoneme features and at least one target evaluation task into a pre-generated pronunciation evaluation model to obtain an evaluation value corresponding to each target evaluation task output by the pronunciation evaluation model. The pronunciation evaluation model is obtained by training a target neural network model through a plurality of sample sets. The sample set includes sample phoneme information of each sample phoneme of a sample text, sample phoneme features of each sample phoneme in sample speech, a plurality of sample evaluation tasks, and sample evaluation values corresponding to each sample evaluation task. The sample phoneme information includes a sample phoneme identifier and sample position information, and the sample position information is used to represent the position of the sample phoneme in a sample phoneme sequence corresponding to the sample text.
9. A computer readable medium having stored thereon a computer program, characterized in that, The program, when executed by the processing device, implements the steps of the method of any one of claims 1-7.
10. An electronic device, comprising: comprising: a storage device having stored thereon at least one computer program; at least one processing device configured to execute the at least one computer program in the storage device to implement the steps of the method of any one of claims 1-7.
Citation Information
Patent Citations
Voice evaluation method and device, storage medium and electronic device
CN110782921A
Online spoken language pronunciation evaluation method and device and storage medium
CN112908360A