Voice quality determination method, apparatus, medium, and electronic device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING YOUZHUJU NETWORK TECH CO LTD
- Filing Date
- 2023-01-17
- Publication Date
- 2026-08-07
AI Technical Summary
[0019]通过上述技术方案,由于音素的实际声学特征与音素的音素标识对应的标准声学特征之间的相似度值表征音素的实际声学特征与音素的音素标识之间的距离,从而增加相似度值有利于区分音素的实际声学特征和音素的音素标识对应的标准声学特征,提高实际声学特征和标准声学特征的特征表达能力;在此基础上,基于音素的实际声学特征、音素的音素标识对应的标准声学特征以及相似度值,确定目标音频的语音质量分数,有利于提高确定的语音质量分数的准确性。
Smart Images

Figure CN116072151B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of speech technology, and more specifically, to a method, apparatus, medium, and electronic device for determining speech quality. Background Technology
[0002] Computer-Aided Pronunciation Training (CAPT) is increasingly used in language teaching, providing support for second language learners to learn a foreign language independently. Pronunciation quality assessment, as a key technology in CAPT, is primarily used to evaluate the accuracy of learners' spoken pronunciation, enabling second language learners to understand their level of mastery of foreign language pronunciation based on accuracy.
[0003] In related technologies, information from both acoustic and text dimensions is used to evaluate the accuracy of learners' spoken pronunciation. However, there is a lack of validation of the effectiveness of combining acoustic and text dimensions, which makes it unclear whether it is appropriate to use text dimension information to evaluate the accuracy of learners' spoken pronunciation. Therefore, how to introduce text dimension information and combine it with acoustic dimension information to evaluate the accuracy of learners' spoken pronunciation is an urgent problem to be solved. Summary of the Invention
[0004] This section is provided to briefly introduce the concepts, which will be described in detail in the Detailed Description section later. This section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] In a first aspect, this disclosure provides a method for determining speech quality, including:
[0006] Obtain the target audio;
[0007] Determine the actual acoustic features of each phoneme in the target text corresponding to the target audio in the target audio, and determine the phoneme identifier of the phoneme;
[0008] The similarity between the actual acoustic features of the phoneme and the standard acoustic features corresponding to the phoneme identifier is calculated to obtain a similarity value;
[0009] The speech quality score of the target audio is determined based on the actual acoustic features of all the phonemes, the standard acoustic features corresponding to the phoneme identifiers of all the phonemes, and all the similarity values, wherein the speech quality score is used to characterize speech quality.
[0010] Secondly, this disclosure provides a method for determining speech quality, including:
[0011] The first acquisition module is used to acquire the target audio;
[0012] The first determining module is used to determine the actual acoustic features of each phoneme in the target text corresponding to the target audio in the target audio, and to determine the phoneme identifier of the phoneme;
[0013] The first calculation module is used to calculate the similarity between the actual acoustic features of the phoneme and the standard acoustic features corresponding to the phoneme identifier, and obtain a similarity value.
[0014] The second determining module is used to determine the speech quality score of the target audio based on the actual acoustic features of all the phonemes, the standard acoustic features corresponding to the phoneme identifiers of all the phonemes, and all the similarity values, wherein the speech quality score is used to characterize speech quality.
[0015] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.
[0016] Fourthly, this disclosure provides an electronic device, comprising:
[0017] A storage device having at least one computer program stored thereon;
[0018] At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method in the first aspect.
[0019] Through the above technical solution, since the similarity value between the actual acoustic features of a phoneme and the standard acoustic features corresponding to the phoneme's phoneme identifier represents the distance between the actual acoustic features of the phoneme and the phoneme's phoneme identifier, increasing the similarity value helps to distinguish between the actual acoustic features of the phoneme and the standard acoustic features corresponding to the phoneme's phoneme identifier, thereby improving the feature representation ability of the actual acoustic features and the standard acoustic features. On this basis, based on the actual acoustic features of the phoneme, the standard acoustic features corresponding to the phoneme's phoneme identifier, and the similarity value, the speech quality score of the target audio is determined, which helps to improve the accuracy of the determined speech quality score.
[0020] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0021] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0022] Figure 1 This is a flowchart illustrating a speech quality determination method according to an exemplary embodiment.
[0023] Figure 2 This is a schematic diagram illustrating a pre-training process according to an exemplary embodiment.
[0024] Figure 3 This is a schematic diagram illustrating a process for determining the speech quality score of a target audio file according to an exemplary embodiment.
[0025] Figure 4 This is a block diagram illustrating a speech quality determination device according to an exemplary embodiment.
[0026] Figure 5 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment. Detailed Implementation
[0027] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0028] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0029] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0031] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0032] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0033] All actions involving the acquisition of signals, information, or data in this disclosure are carried out in accordance with the relevant data protection laws and policies of the country where the location is situated, and with the authorization granted by the owner of the relevant device.
[0034] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0035] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0036] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0037] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0038] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0039] As mentioned in the background section, related technologies use models to process acoustic and textual information to evaluate the accuracy of second language learners' spoken pronunciation. Here, "second language learner" refers to a learner whose first language is their native language and who is learning a second language. However, these technologies typically combine acoustic and textual information by simply summing or concatenating features. There's no clear concept of this combination; for example, it doesn't specify whether a larger distance between a phoneme-level acoustic feature and its corresponding phoneme identifier indicates poor pronunciation, or a smaller distance indicates good pronunciation. This fails to explain the effectiveness of the combination method, leading to uncertainty about whether incorporating textual information to evaluate the accuracy of second language learners' spoken pronunciation is appropriate. Therefore, how to incorporate textual information and combine it with acoustic information to evaluate the accuracy of learners' spoken pronunciation is a pressing issue.
[0040] In view of this, embodiments of the present disclosure provide a method, apparatus, medium, and electronic device for determining speech quality, which improves the accuracy of the determined speech quality score.
[0041] The embodiments of this disclosure will be further explained below with reference to the accompanying drawings.
[0042] Figure 1 This is a flowchart illustrating a speech quality determination method according to an exemplary embodiment. This speech quality determination method can be applied to electronic devices, such as mobile terminals (e.g., mobile phones, tablets), and fixed terminals (e.g., servers, desktop computers). (Refer to...) Figure 1 The speech quality determination method may include the following steps:
[0043] Step S101: Obtain the target audio.
[0044] The target audio is audio that includes the user's speech data, where the user is similar to the second language learner mentioned above. For example, the speech data could be the English speech of a user learning English, or, as another example, the Chinese speech of a user learning Chinese.
[0045] For example, the target audio can be audio recorded by an electronic device and stored locally, or audio obtained from a remote server.
[0046] Step S102: Determine the actual acoustic features of each phoneme in the target text corresponding to the target audio in the target audio, and determine the phoneme identifier of the phoneme.
[0047] Among them, the target audio can be converted into the corresponding target text through an ASR (Automatic Speech Recognition) model. For the specific implementation of the ASR model, reference can be made to related technologies, which will not be elaborated in this embodiment.
[0048] Among them, a phoneme is the smallest speech unit divided according to the natural attributes of speech. For example, the Chinese character "啊" has only one phoneme "a", "爱" has two phonemes "a" and "i", and "代" has three phonemes "d", "a", and "i".
[0049] Among them, the actual acoustic features of a phoneme are used to represent the actual pronunciation of the phoneme by the user. The actual acoustic features can be extracted based on the MFCC (Mel-Frequency Cepstrum Coefficient) features of the target audio. In sound processing, MFC (Mel-Frequency Cepstrum) represents the short-time power spectrum of sound. MFC is a linear cosine transform of the logarithmic power spectrum based on the non-linear Mel scale frequency. The MFCC features are all the coefficients that make up the MFC. For example, the MFCC features can be obtained in the following way: after performing operations such as digital-to-analog conversion, pre-emphasis, frame addition with windowing, and discrete Fourier transform on the target audio, the MFCC features can be obtained. In addition, for the specific implementation process of extracting the actual acoustic features of phonemes in the target audio based on the MFCC features, reference can be made to the following related embodiments, which will not be elaborated in this embodiment.
[0050] Among them, a phoneme identifier is used to uniquely correspond to a phoneme. Through this phoneme identifier, the corresponding standard acoustic features can be obtained. Among them, the standard acoustic features of the phoneme identifier are used to represent the standard pronunciation of the phoneme corresponding to the phoneme identifier. It should be noted that the phoneme identifiers of each phoneme can be constructed in advance. In addition, for the specific implementation process of determining the phoneme identifier, reference can be made to the following related embodiments, which will not be elaborated in this embodiment.
[0051] Step S103, calculate the similarity between the actual acoustic features of the phoneme and the standard acoustic features corresponding to the phoneme identifier of the phoneme, and obtain a similarity value.
[0052] Among them, the similarity value is proportional to the speech quality. The higher the similarity value, the closer the actual acoustic features and the standard acoustic features are, and the higher the pronunciation quality, that is, the higher the speech quality. For example, the similarity calculation can be a cosine distance calculation, and then the similarity value can be a cosine distance, and the smaller the cosine distance, the higher the similarity value.
[0053] Step S104: Determine the speech quality score of the target audio based on the actual acoustic features of all phonemes, the standard acoustic features corresponding to the phoneme identifiers of all phonemes, and all similarity values. The speech quality score is used to characterize speech quality.
[0054] Among them, the speech quality score is directly proportional to the speech quality; the higher the speech quality score, the higher the speech quality.
[0055] It is worth noting that the actual acoustic features of a phoneme, the standard acoustic features corresponding to the phoneme's phoneme identifier, and the similarity value are all phoneme-level features. The speech quality score of the target audio is determined by comprehensively considering the actual acoustic features, standard acoustic features, and similarity value at the phoneme level. Furthermore, the specific implementation process for determining the speech quality score of the target audio can be found in the following related embodiments, which will not be elaborated upon here.
[0056] Through the above technical solution, since the similarity value between the actual acoustic features of a phoneme and the standard acoustic features corresponding to the phoneme's phoneme identifier represents the distance between the actual acoustic features of the phoneme and the phoneme's phoneme identifier, increasing the similarity value helps to distinguish between the actual acoustic features of the phoneme and the standard acoustic features corresponding to the phoneme's phoneme identifier, thereby improving the feature representation ability of the actual acoustic features and the standard acoustic features. On this basis, based on the actual acoustic features of the phoneme, the standard acoustic features corresponding to the phoneme's phoneme identifier, and the similarity value, the speech quality score of the target audio is determined, which helps to improve the accuracy of the determined speech quality score.
[0057] In some embodiments, the actual acoustic features of phonemes and the phoneme identifiers can be determined using a pre-trained acoustic model. For example, step S102 above can be implemented as follows: the acoustic features corresponding to the target audio are processed by the pre-trained acoustic model to obtain the phoneme likelihood value corresponding to each speech frame in the target audio and the target features output by the target layer of the acoustic model; the phoneme timestamp of each phoneme in the target text is determined based on the target text corresponding to the target audio and the phoneme likelihood value; the actual acoustic features of the phoneme in the target audio are determined based on the target features corresponding to the speech frame of each phoneme's phoneme timestamp.
[0058] The acoustic model can be a DNN (Deep Neural Network)-Hidden Markov Model (DNN-HMM). The training of the acoustic model can be referred to in the following related embodiments.
[0059] In this embodiment, the acoustic features corresponding to the target audio can be the MFCC features mentioned above.
[0060] The phoneme likelihood value is used to characterize the probability of a speech frame corresponding to each phoneme, and it can also be used to indicate the phoneme corresponding to a speech frame. It is understood that the phoneme with the highest probability is the phoneme corresponding to the speech frame.
[0061] The target layer of the acoustic model can be the bottleneck layer of the acoustic model.
[0062] The encoder can be used to process the phoneme likelihood values and the target text to obtain the phoneme timestamps. Each phoneme timestamp includes a phoneme identifier and one or more corresponding speech frames. For example, taking the target text "good" as an example, the phonemes corresponding to "good" are "g", "ow", and "d". Specifically, in the target audio, the first and second frames all contain the phoneme "g", meaning the phoneme timestamp for "g" includes the phoneme identifier corresponding to "g" and the first and second frames corresponding to "g"; the third to fifth frames all contain the phoneme "ow", meaning the phoneme timestamp for "ow" includes the phoneme identifier corresponding to "ow" and the third to fifth frames corresponding to "ow"; and the sixth to tenth frames all contain the phoneme "d", meaning the phoneme timestamp for "d" includes the phoneme identifier corresponding to "d" and the sixth to tenth frames corresponding to "d".
[0063] Understandably, after determining the phoneme timestamp, the target features corresponding to each speech frame within that phoneme timestamp can be found. Then, based on the target features corresponding to each speech frame within that phoneme timestamp, the actual acoustic features at the phoneme level can be determined. For example, the average feature of the target features corresponding to the speech frames at the phoneme timestamp can be determined as the actual acoustic feature of that phoneme in the target audio. For instance, continuing the example above, for the phoneme "g", the target feature corresponding to the first frame is K1, and the target feature corresponding to the second frame is K2. Then, the actual acoustic feature of the phoneme "g" is K3, where K3 is the average of K1 and K2.
[0064] Using the above method, the actual acoustic characteristics at the phoneme level can be determined based on the acoustic model.
[0065] Furthermore, based on the aforementioned phoneme likelihood values and the target text corresponding to the target audio, phonemes can be determined, thereby enabling the identification of phonemes. The phoneme identification can be carried within the aforementioned phoneme timestamps.
[0066] In some embodiments, the acoustic model described above can be trained by: obtaining a first training speech sample set; and training an initial acoustic model based on the first training samples in the first training sample set to obtain a trained acoustic model.
[0067] The first training sample set can be constructed by collecting approximately 960 hours of audio data from native English speakers and approximately 10 hours of audio data from native Chinese speakers learning English. The audio data from native English speakers is used to construct standard acoustic features, while the audio data from native Chinese speakers learning English is used to construct actual acoustic features.
[0068] For example, the first training sample includes sample MFCC features and corresponding language frame labels, where the speech frame labels represent the corresponding sample phonemes. A DFSMN (Deep Feedforward Sequential Memory Networks)-Hidden Markov Model is trained using the sample MFCC features and corresponding language frame labels. The trained DFSMN-Hidden Markov Model can then serve as a trained acoustic model, used to determine the phoneme likelihood value for each speech frame and the target features output by the target layer of the acoustic model. It is worth noting that the DFSMN-Hidden Markov Model serves as the initial acoustic model, while DFSMN can be considered a type of DNN network, and other DNN networks can also be used to construct the network structure of the initial acoustic model.
[0069] It is worth noting that the first training speech sample can be constructed in the following way: First, extract the sample MFCC features of the sample audio, train the Gaussian Mixture Model-Hidden Markov Model (GMM-HMM) using the sample MFCC features, and perform audio alignment operation with the sample text corresponding to the sample audio to obtain speech frame labels. Thus, the first training speech sample is constructed based on the sample MFCC features and the corresponding speech frame labels, which is used to train the DFSMN-Hidden Markov Model.
[0070] In some embodiments, similarity values and speech quality scores can be determined using a pre-trained speech quality model. For example, step S103 above can be implemented as follows: extract feature vectors corresponding to the actual acoustic features of phonemes and feature vectors corresponding to phoneme identifiers using the pre-trained speech quality model; calculate the similarity of the extracted feature vectors using the speech quality model to obtain a similarity value; simultaneously, step S104 above can be implemented as follows: process the actual acoustic features of all phonemes, the standard acoustic features corresponding to the phoneme identifiers of all phonemes, and all similarity values using the speech quality model to obtain the speech quality score of the target audio.
[0071] As an example, a speech quality model may include a first feature vector extraction layer and a second feature vector extraction layer. The first feature vector extraction layer can be used to extract feature vectors corresponding to the actual acoustic features of phonemes; the second feature vector extraction layer can be used to extract feature vectors corresponding to phoneme identifiers as the standard acoustic features of the phoneme identifiers. It is worth noting that the first feature vector extraction layer is used to reduce the dimensionality of the actual acoustic features of phonemes, thereby obtaining feature vectors of the same dimension as the standard acoustic features of phoneme identifiers, which facilitates similarity calculation.
[0072] As an example, by calculating the cosine distance between the feature vectors extracted by the first feature vector extraction layer and the feature vectors extracted by the second feature vector extraction layer using the speech quality model, the similarity value between the actual acoustic features of the phoneme and the standard acoustic features corresponding to the phoneme identifier can be obtained.
[0073] As an example, the speech quality model may also include a bidirectional LSTM (Long Short-Term Memory) network. This network processes the actual acoustic features of all phonemes, the standard acoustic features corresponding to the phoneme identifiers of all phonemes, and all similarity values to obtain the speech quality score of the target audio. The specific model structure of the bidirectional LSTM network can be found in relevant technologies, and will not be elaborated upon in this embodiment.
[0074] As an example, step S104 above can further include: processing the feature vectors corresponding to the actual acoustic features of all phonemes, the standard acoustic features corresponding to the phoneme identifiers of all phonemes, and all similarity values through a speech quality model to obtain the speech quality score of the target audio. Here, the feature vectors corresponding to the actual acoustic features are feature vectors obtained by dimensionality reduction processing of the actual acoustic features through the first feature vector extraction layer.
[0075] As an example, the speech quality score of the target audio is obtained by processing the feature vectors corresponding to the actual acoustic features of all phonemes, the standard acoustic features corresponding to the phoneme identifiers of all phonemes, and all similarity values through a bidirectional LSTM network in the speech quality model.
[0076] In some embodiments, the speech quality model can be trained by: acquiring a second training speech sample set and a third training speech sample set; pre-training a network layer that extracts the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifier using the second training speech samples in the second training speech sample set; and training the initial speech quality model using the third training speech samples in the third training speech sample set to obtain a trained speech quality model, wherein the initial speech quality model includes a pre-trained network layer for extracting the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifier.
[0077] The specific implementation process of pre-training the network layer that extracts the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifier can be referred to in the following relevant embodiments, which will not be elaborated here.
[0078] The third training speech sample includes sample phoneme identifiers, actual acoustic features of the sample, and manually labeled sample speech quality scores. The sample speech quality scores are used as supervision information to train the initial speech quality model.
[0079] The third training speech sample set was constructed by randomly collecting approximately 2,500 audio recordings of second language learners' pronunciations and their accompanying texts. Language experts scored each sentence according to accuracy scoring rules. Here, the scoring refers to the manually labeled quality score of the sample speech.
[0080] It is worth noting that the network layer used to extract the first sample feature vector corresponding to the actual acoustic features of the sample phonemes is equivalent to the function of the first feature vector extraction layer mentioned above, and the network layer used to extract the second sample feature vector corresponding to the sample phoneme identifier is equivalent to the function of the second feature vector extraction layer mentioned above. In other words, the trained network layer used to extract the first sample feature vector corresponding to the actual acoustic features of the sample phonemes is equivalent to the first feature vector extraction layer mentioned above; and the trained network layer used to extract the second sample feature vector corresponding to the sample phoneme identifier is equivalent to the second feature vector extraction layer mentioned above.
[0081] It is worth noting that the initial speech quality model, in addition to the network layers that extract the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers, can also include a network similar to a bidirectional LSTM. It is also worth noting that the network parameters of the network layers mentioned here (including the network layers that extract the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers) and the network (i.e., the network similar to a bidirectional LSTM) are fine-tuned using a third training speech sample, thereby obtaining a trained speech quality model.
[0082] The above method firstly utilizes the second training speech sample set to pre-train the network layer that extracts the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers, thereby improving the network layer's feature representation ability for the first and second sample feature vectors. Consequently, when training the initial speech quality model based on the third training speech sample set, the correlation with manually labeled sample speech quality score labels is improved, thus enhancing the model's performance.
[0083] In some embodiments, the network layers that extract the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers can be pre-trained in the following manner: extracting the first sample feature vector corresponding to the actual acoustic features of the samples in the second training speech samples in the second training speech sample set, and extracting the second sample feature vector corresponding to the sample phoneme identifiers in the second training speech samples; calculating the similarity between the first sample feature vector and the second sample feature vector to obtain a sample similarity value; and adjusting the network parameters of the network layers that extract the first sample feature vector and the second sample feature vector respectively based on the difference between the sample similarity and the pronunciation goodness of the samples in the second training speech samples.
[0084] The second training speech sample set includes sample phoneme identifiers, sample actual acoustic features, and sample goodness of pronunciation (GOP). The GOP serves as supervisory information, using phoneme-level sample phoneme identifiers and phoneme-level sample actual acoustic features to optimize the network parameters of the network layers that extract the feature vectors of the first and second samples, respectively. For example, the MSE loss function can be used to characterize the difference between the first and second sample feature vectors.
[0085] It is worth noting that the first sample feature vector corresponding to the actual acoustic features of the sample here is the feature vector obtained by dimensionality reduction of the actual acoustic features of the sample.
[0086] Figure 2 This is a schematic diagram of pre-training according to an exemplary embodiment. The pre-training here refers to pre-training a network layer that extracts a first sample feature vector corresponding to the actual acoustic features of the sample phonemes and a second sample feature vector corresponding to the sample phoneme identifier. (Refer to...) Figure 2 , Figure 2 The linear layer shown is the network layer that extracts the feature vector of the first sample. Figure 2 The Embedding layer shown is the network layer for extracting the feature vector of the second sample. The sample phoneme identifier is processed by the Embedding layer to obtain the second sample feature vector; the actual acoustic features of the sample are processed by the linear layer to obtain the first sample feature vector; the cosine distance between the first sample feature vector and the second sample feature vector is calculated to obtain the sample similarity value; the network parameters of the Embedding layer and the linear layer are adjusted based on the sample similarity value and the sample pronunciation goodness.
[0087] It is worth noting that the linear layer here is used to reduce the dimensionality of the actual acoustic features of the sample, so as to obtain the first sample feature vector with the same dimension as the second sample feature vector.
[0088] By using the above method, the network parameters of the network layer used to extract the first sample feature vector and the second sample feature vector are pre-trained using the phoneme-level pronunciation goodness, thereby improving the representation ability of the first sample feature vector and the second sample feature vector and thus improving the performance of the model.
[0089] Figure 3 This is a schematic diagram illustrating a process for determining the speech quality score of target audio according to an exemplary embodiment. (Refer to...) Figure 3First, the input audio is acquired, and features are extracted to obtain MFCC features. The MFCC features are then input into an acoustic model to obtain the phoneme likelihood value and frame-level bottleneck features for each speech frame in the input audio. The phoneme likelihood value for each speech frame and the corresponding input text are input into a decoder to obtain the phoneme timestamp for each phoneme output by the decoder. This timestamp includes the phoneme ID and one or more corresponding speech frames. Using the phoneme timestamp and frame-level bottleneck features, phoneme-level frame bottleneck features are determined. The phoneme ID and frame-level bottleneck features are then input into a speech quality model. The embedding layer in the speech quality model extracts the feature vector of the phoneme ID, and the linear layer extracts the feature vector of the phoneme-level frame bottleneck features. The cosine distance between the feature vectors of the phoneme ID and the phoneme-level frame bottleneck features is calculated. Finally, the feature vectors of the phoneme ID, the phoneme-level frame bottleneck features, and the cosine distance are input into a BLSTM (Bidirectional Long Short-Term Memory) system. The accuracy score is obtained by using a bidirectional long short-term memory (BSSM) scoring model.
[0090] It is worth noting that the input audio here is, for example, the target audio mentioned above; the frame-level bottleneck features here are, for example, the target features mentioned above; the input text here is, for example, the target text mentioned above, which corresponds to the input audio; the phoneme-level frame bottleneck features here are, for example, the actual acoustic features mentioned above; the BLSTM scoring model here is, for example, the bidirectional LSTM network mentioned above; and the accuracy scoring result here is, for example, the speech quality score mentioned above.
[0091] Figure 4 This is a block diagram illustrating a voice quality determination apparatus according to an exemplary embodiment. (Refer to...) Figure 4 The voice quality determination device includes:
[0092] The first acquisition module 401 is used to acquire the target audio;
[0093] The first determining module 402 is used to determine the actual acoustic features of each phoneme in the target text corresponding to the target audio in the target audio, and to determine the phoneme identifier of the phoneme;
[0094] The first calculation module 403 is used to calculate the similarity between the actual acoustic features of the phoneme and the standard acoustic features corresponding to the phoneme identifier, and obtain a similarity value.
[0095] The second determining module 404 is used to determine the speech quality score of the target audio based on the actual acoustic features of all the phonemes, the standard acoustic features corresponding to the phoneme identifiers of all the phonemes, and all the similarity values, wherein the speech quality score is used to characterize speech quality.
[0096] Optionally, the first determining module 402 includes:
[0097] The first processing submodule is used to process the acoustic features corresponding to the target audio through a pre-trained acoustic model to obtain the phoneme likelihood value corresponding to each speech frame in the target audio and the target features output by the target layer of the acoustic model.
[0098] The first determining submodule is used to determine the phoneme timestamp of each phoneme in the target text based on the target text corresponding to the target audio and the phoneme likelihood value;
[0099] The second determining submodule is used to determine the actual acoustic features of the phoneme in the target audio based on the target features corresponding to the speech frame corresponding to the phoneme timestamp of each phoneme.
[0100] Optionally, the second determining submodule is specifically used to: determine the average feature of the target feature corresponding to the speech frame corresponding to the phoneme timestamp of each phoneme as the actual acoustic feature of the phoneme in the target audio.
[0101] Optionally, the first computing module 403 includes:
[0102] The extraction submodule is used to extract the feature vectors corresponding to the actual acoustic features of the phonemes and the feature vectors corresponding to the phoneme identifiers using a pre-trained speech quality model.
[0103] The calculation submodule is used to calculate the similarity of the extracted feature vectors using the speech quality model to obtain a similarity value;
[0104] The second determining module 404 is used to process the actual acoustic features of all the phonemes, the standard acoustic features corresponding to the phoneme identifiers of all the phonemes, and all the similarity values through the speech quality model to obtain the speech quality score of the target audio.
[0105] Optionally, the speech quality determination device 400 further includes a first training module, through which the acoustic model is trained. The first training module includes:
[0106] The first acquisition submodule is used to acquire the first training speech sample set;
[0107] The first training submodule is used to train an initial acoustic model based on the first training samples in the first training sample set to obtain a trained acoustic model.
[0108] Optionally, the speech quality determination device 400 further includes a second training module, through which the speech quality model is trained. The second training module includes:
[0109] The second acquisition submodule is used to acquire the second training speech sample set and the third training speech sample set;
[0110] The second training submodule is used to pre-train the network layer that extracts the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers using the second training speech samples in the second training speech sample set.
[0111] The third training submodule is used to train the initial speech quality model using the third training speech samples in the third training speech sample set to obtain a trained speech quality model. The initial speech quality model includes a pre-trained network layer for extracting the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers.
[0112] Optionally, the second training submodule is specifically used for:
[0113] Extract the first sample feature vector corresponding to the actual acoustic features of the samples in the second training speech sample set, and extract the second sample feature vector corresponding to the phoneme identifiers of the samples in the second training speech sample set;
[0114] The similarity between the first sample feature vector and the second sample feature vector is calculated to obtain the sample similarity value;
[0115] Based on the difference between the sample similarity and the pronunciation quality of the samples in the second training speech samples, the network parameters of the network layers that extract the feature vectors of the first sample and the feature vectors of the second sample are adjusted.
[0116] The specific implementation methods of each module of the above-mentioned device can be referred to the above-mentioned related embodiments, and will not be repeated here.
[0117] This disclosure also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the above method embodiments.
[0118] This disclosure also provides an electronic device, comprising:
[0119] A storage device having at least one computer program stored thereon;
[0120] At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method described in the above method embodiments.
[0121] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device 500 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0122] like Figure 5 As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0123] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0124] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0125] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0126] In some implementations, electronic devices can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0127] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0128] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire target audio; determine the actual acoustic features of each phoneme in the target audio within the target audio, and determine the phoneme identifier; perform similarity calculation on the actual acoustic features of the phoneme and the standard acoustic features corresponding to the phoneme identifier to obtain a similarity value; and determine the speech quality score of the target audio based on the actual acoustic features of all the phonemes, the standard acoustic features corresponding to the phoneme identifiers of all the phonemes, and all the similarity values, wherein the speech quality score is used to characterize speech quality.
[0129] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0130] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0131] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the functionality of the module itself.
[0132] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0133] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0134] According to one or more embodiments of this disclosure, Example 1 provides a method for determining speech quality, including:
[0135] Obtain the target audio;
[0136] Determine the actual acoustic features of each phoneme in the target text corresponding to the target audio in the target audio, and determine the phoneme identifier of the phoneme;
[0137] The similarity between the actual acoustic features of the phoneme and the standard acoustic features corresponding to the phoneme identifier is calculated to obtain a similarity value;
[0138] The speech quality score of the target audio is determined based on the actual acoustic features of all the phonemes, the standard acoustic features corresponding to the phoneme identifiers of all the phonemes, and all the similarity values, wherein the speech quality score is used to characterize speech quality.
[0139] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein determining the actual acoustic features of each phoneme in the target text corresponding to the target audio in the target audio includes:
[0140] The acoustic features corresponding to the target audio are processed by a pre-trained acoustic model to obtain the phoneme likelihood value corresponding to each speech frame in the target audio and the target features output by the target layer of the acoustic model.
[0141] Based on the target text corresponding to the target audio and the phoneme likelihood value, determine the phoneme timestamp of each phoneme in the target text;
[0142] Based on the target features corresponding to the speech frame corresponding to the phoneme timestamp of each phoneme, the actual acoustic features of the phoneme in the target audio are determined.
[0143] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 1, wherein determining the actual acoustic features of a phoneme in the target audio based on the target features corresponding to the speech frame corresponding to the phoneme timestamp of each phoneme includes:
[0144] The average feature of the target feature corresponding to the speech frame corresponding to the phoneme timestamp of each phoneme is determined as the actual acoustic feature of that phoneme in the target audio.
[0145] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 1, wherein calculating the similarity between the actual acoustic features of the phoneme and the standard acoustic features corresponding to the phoneme identifier of the phoneme to obtain a similarity value includes:
[0146] The feature vectors corresponding to the actual acoustic features of the phonemes and the feature vectors corresponding to the phoneme identifiers are extracted using a pre-trained speech quality model.
[0147] The similarity of the extracted feature vectors is calculated using the aforementioned speech quality model to obtain a similarity value.
[0148] The step of determining the speech quality score of the target audio based on the actual acoustic features of all the phonemes, the standard acoustic features corresponding to the phoneme identifiers of all the phonemes, and all the similarity values includes:
[0149] The speech quality model is used to process the actual acoustic features of all the phonemes, the standard acoustic features corresponding to the phoneme identifiers of all the phonemes, and all the similarity values to obtain the speech quality score of the target audio.
[0150] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 1, wherein the acoustic model is trained in the following manner:
[0151] Obtain the first training speech sample set;
[0152] The initial acoustic model is trained based on the first training sample in the first training sample set to obtain the trained acoustic model.
[0153] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 1, wherein the speech quality model is trained in the following manner:
[0154] Obtain the second and third training speech sample sets;
[0155] The network layers that extract the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers are pre-trained using the second training speech samples in the second training speech sample set.
[0156] The initial speech quality model is trained using the third training speech samples in the third training speech sample set to obtain a trained speech quality model. The initial speech quality model includes a pre-trained network layer for extracting the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers.
[0157] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 1, wherein the network layer that extracts a first sample feature vector corresponding to the actual acoustic features of the sample phonemes and a second sample feature vector corresponding to the sample phoneme identifiers respectively from the second training speech samples in the second training speech sample set is pre-trained, including:
[0158] Extract the first sample feature vector corresponding to the actual acoustic features of the samples in the second training speech sample set, and extract the second sample feature vector corresponding to the phoneme identifiers of the samples in the second training speech sample set;
[0159] The similarity between the first sample feature vector and the second sample feature vector is calculated to obtain the sample similarity value;
[0160] Based on the difference between the sample similarity and the pronunciation quality of the samples in the second training speech samples, the network parameters of the network layers that extract the feature vectors of the first sample and the feature vectors of the second sample are adjusted.
[0161] According to one or more embodiments of this disclosure, Example 8 provides a speech quality determination apparatus, comprising:
[0162] The first acquisition module is used to acquire the target audio;
[0163] The first determining module is used to determine the actual acoustic features of each phoneme in the target text corresponding to the target audio in the target audio, and to determine the phoneme identifier of the phoneme;
[0164] The first calculation module is used to calculate the similarity between the actual acoustic features of the phoneme and the standard acoustic features corresponding to the phoneme identifier, and obtain a similarity value.
[0165] The second determining module is used to determine the speech quality score of the target audio based on the actual acoustic features of all the phonemes, the standard acoustic features corresponding to the phoneme identifiers of all the phonemes, and all the similarity values, wherein the speech quality score is used to characterize speech quality.
[0166] According to one or more embodiments of this disclosure, Example 9 provides the apparatus of Example 8, wherein the first determining module includes:
[0167] The first processing submodule is used to process the acoustic features corresponding to the target audio through a pre-trained acoustic model to obtain the phoneme likelihood value corresponding to each speech frame in the target audio and the target features output by the target layer of the acoustic model.
[0168] The first determining submodule is used to determine the phoneme timestamp of each phoneme in the target text based on the target text corresponding to the target audio and the phoneme likelihood value;
[0169] The second determining submodule is used to determine the actual acoustic features of the phoneme in the target audio based on the target features corresponding to the speech frame corresponding to the phoneme timestamp of each phoneme.
[0170] According to one or more embodiments of this disclosure, Example 10 provides the apparatus of Example 9, wherein the second determining submodule is specifically configured to: determine the average feature of the target feature corresponding to the speech frame corresponding to the phoneme timestamp of each phoneme as the actual acoustic feature of the phoneme in the target audio.
[0171] According to one or more embodiments of this disclosure, Example 11 provides an apparatus of Example 9, wherein the first computing module includes:
[0172] The extraction submodule is used to extract the feature vectors corresponding to the actual acoustic features of the phonemes and the feature vectors corresponding to the phoneme identifiers using a pre-trained speech quality model.
[0173] The calculation submodule is used to calculate the similarity of the extracted feature vectors using the speech quality model to obtain a similarity value;
[0174] The second determining module is used to process the actual acoustic features of all the phonemes, the standard acoustic features corresponding to the phoneme identifiers of all the phonemes, and all the similarity values through the speech quality model to obtain the speech quality score of the target audio.
[0175] According to one or more embodiments of this disclosure, Example 12 provides the apparatus of Example 9, wherein the speech quality determination apparatus further includes a first training module through which the acoustic model is trained, the first training module comprising:
[0176] The first acquisition submodule is used to acquire the first training speech sample set;
[0177] The first training submodule is used to train an initial acoustic model based on the first training samples in the first training sample set to obtain a trained acoustic model.
[0178] According to one or more embodiments of this disclosure, Example 13 provides the apparatus of Example 11, wherein the speech quality determination apparatus further includes a second training module, through which the speech quality model is trained, the second training module comprising:
[0179] The second acquisition submodule is used to acquire the second training speech sample set and the third training speech sample set;
[0180] The second training submodule is used to pre-train the network layer that extracts the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers using the second training speech samples in the second training speech sample set.
[0181] The third training submodule is used to train the initial speech quality model using the third training speech samples in the third training speech sample set to obtain a trained speech quality model. The initial speech quality model includes a pre-trained network layer for extracting the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers.
[0182] According to one or more embodiments of this disclosure, Example 14 provides the apparatus of Example 13, wherein the second training submodule is specifically used for:
[0183] Extract the first sample feature vector corresponding to the actual acoustic features of the samples in the second training speech sample set, and extract the second sample feature vector corresponding to the phoneme identifiers of the samples in the second training speech sample set;
[0184] The similarity between the first sample feature vector and the second sample feature vector is calculated to obtain the sample similarity value;
[0185] Based on the difference between the sample similarity and the pronunciation quality of the samples in the second training speech samples, the network parameters of the network layers that extract the feature vectors of the first sample and the feature vectors of the second sample are adjusted.
[0186] According to one or more embodiments of the present disclosure, Example 15 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any one of Examples 1-7.
[0187] According to one or more embodiments of this disclosure, Example 16 provides an electronic device comprising:
[0188] A storage device having at least one computer program stored thereon;
[0189] At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of any one of the methods described in Examples 1-7.
[0190] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0191] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0192] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. A method for determining speech quality, characterized in that, include: Obtain the target audio; Determine the actual acoustic features of each phoneme in the target text corresponding to the target audio in the target audio, and determine the phoneme identifier of the phoneme; The similarity value is obtained by calculating the similarity between the actual acoustic features of the phoneme and the standard acoustic features corresponding to the phoneme identifier using a pre-trained speech quality model. The speech quality model determines the speech quality score of the target audio based on the actual acoustic features of all phonemes, the standard acoustic features corresponding to the phoneme identifiers of all phonemes, and all similarity values. The speech quality score characterizes the speech quality. The speech quality model is trained as follows: a second training speech sample set and a third training speech sample set are obtained; a network layer is pre-trained using the second training speech samples from the second training speech sample set to extract the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers; and an initial speech quality model is trained using the third training speech samples from the third training speech sample set to obtain the speech quality model. The initial speech quality model includes pre-trained network layers for extracting the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers.
2. The method according to claim 1, characterized in that, Determining the actual acoustic features of each phoneme in the target text corresponding to the target audio in the target audio includes: The acoustic features corresponding to the target audio are processed by a pre-trained acoustic model to obtain the phoneme likelihood value corresponding to each speech frame in the target audio and the target features output by the target layer of the acoustic model. Based on the target text corresponding to the target audio and the phoneme likelihood value, determine the phoneme timestamp of each phoneme in the target text; Based on the target features corresponding to the speech frame corresponding to the phoneme timestamp of each phoneme, the actual acoustic features of the phoneme in the target audio are determined.
3. The method according to claim 2, characterized in that, The step of determining the actual acoustic features of a phoneme in the target audio based on the target features corresponding to the speech frame corresponding to the phoneme timestamp of each phoneme includes: The average feature of the target feature corresponding to the speech frame corresponding to the phoneme timestamp of each phoneme is determined as the actual acoustic feature of that phoneme in the target audio.
4. The method according to claim 2, characterized in that, The step involves calculating the similarity between the actual acoustic features of the phoneme and the standard acoustic features corresponding to the phoneme identifier using a pre-trained speech quality model to obtain a similarity value, including: The feature vectors corresponding to the actual acoustic features of the phonemes and the feature vectors corresponding to the phoneme identifiers are extracted using a pre-trained speech quality model. The similarity of the extracted feature vectors is calculated using the aforementioned speech quality model to obtain a similarity value.
5. The method according to claim 2, characterized in that, The acoustic model was trained in the following manner: Obtain the first training speech sample set; The initial acoustic model is trained based on the first training sample in the first training sample set to obtain the trained acoustic model.
6. The method according to claim 1, characterized in that, The pre-training of the network layer that extracts the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers from the second training speech samples in the second training speech sample set includes: Extract the first sample feature vector corresponding to the actual acoustic features of the samples in the second training speech sample set, and extract the second sample feature vector corresponding to the phoneme identifiers of the samples in the second training speech sample set; The similarity between the first sample feature vector and the second sample feature vector is calculated to obtain the sample similarity value; Based on the difference between the sample similarity and the pronunciation quality of the samples in the second training speech samples, the network parameters of the network layers that extract the feature vectors of the first sample and the feature vectors of the second sample are adjusted.
7. A voice quality determination device, characterized in that, include: The first acquisition module is used to acquire the target audio; The first determining module is used to determine the actual acoustic features of each phoneme in the target text corresponding to the target audio in the target audio, and to determine the phoneme identifier of the phoneme; The first calculation module is used to calculate the similarity between the actual acoustic features of the phoneme and the standard acoustic features corresponding to the phoneme identifier using a pre-trained speech quality model, and obtain a similarity value. The second determining module is used to determine the speech quality score of the target audio based on the actual acoustic features of all the phonemes, the standard acoustic features corresponding to the phoneme identifiers of all the phonemes, and all the similarity values using the speech quality model. The speech quality score is used to characterize speech quality. The speech quality model is trained as follows: acquiring a second training speech sample set and a third training speech sample set; pre-training a network layer that extracts the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers using the second training speech samples from the second training speech sample set; and training an initial speech quality model using the third training speech samples from the third training speech sample set to obtain the speech quality model. The initial speech quality model includes a pre-trained network layer that extracts the first sample feature vector corresponding to the actual acoustic features of the sample phonemes and the second sample feature vector corresponding to the sample phoneme identifiers.
8. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method according to any one of claims 1-6.
9. An electronic device, characterized in that, include: A storage device having at least one computer program stored thereon; At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method according to any one of claims 1-6.
Citation Information
Patent Citations
Audio frequency processing method, device and system, storage medium, terminal and server
CN110148427A
Method and device for processing voice data, equipment and storage medium
CN115273897A