Method for processing voice data, method for generating model, device, and electronic device

By determining the acoustic characteristics and phoneme data in the speech data and combining the augmented sample speech data training model, the problem of poor robustness of the traditional pronunciation quality evaluation model is solved, and more accurate speech quality evaluation and user experience are achieved.

CN115359808BActive Publication Date: 2025-08-05BEIJING YOUZHUJU NETWORK TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211006080.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-22
Publication Date
2025-08-05
Estimated Expiration
2042-08-22

AI Technical Summary

Technical Problem

The traditional pronunciation quality evaluation model is poorly robust, resulting in poor user experience, and the scarcity of second-language pronunciation corpus and high labeling cost, limiting the performance of pronunciation quality evaluation system.

Method used

By determining the acoustic features in the speech data, extracting frame-level feature data and phoneme data, determining the quality level of speech data based on these data, and using the augmented sample speech data to train the model to improve the evaluation accuracy.

Benefits of technology

The frame-level voice quality evaluation is achieved, which improves the accuracy of voice quality evaluation and the robustness of the system, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359808B_ABST
    Figure CN115359808B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method, a model generation method, an apparatus, and an electronic device for processing speech data. The method may include determining acoustic features corresponding to speech frames in the speech data. The method may also include extracting feature data corresponding to the speech frames from the acoustic features. In addition, the method may further include determining phoneme data corresponding to the speech frames based at least on the acoustic features. The method may also include determining a quality level of the speech data based on the phoneme data and the feature data, wherein the quality level indicates the speech quality of the speech data. The present disclosure implements frame-level speech quality assessment, thereby optimizing the quality level determination process and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of data processing, and more particularly, to a method for processing speech data, a model generation method, an apparatus, an electronic device, and a computer program product. Background Art

[0002] For the pronunciation of various languages, it is very important for users to know whether their pronunciation when reading aloud or following along is standard. With the increasing popularity of online interactive technology, Computer-Aided Pronunciation Training (CAPT) is increasingly being used in users' pronunciation attempts. Pronunciation quality assessment, as an important technology in computer-aided pronunciation training, is mainly used to evaluate the accuracy of users' spoken pronunciation. In the process of technical practice, people have found that traditional pronunciation quality assessment models or applications have poor robustness problems, making the user experience urgently need to be improved. Summary of the Invention

[0003] The embodiments of the present disclosure provide a solution for processing speech data and a solution for generating a model.

[0004] In a first aspect of the present disclosure, a method for processing speech data is provided. The method may include determining acoustic features corresponding to speech frames in the speech data. The method may also include extracting feature data corresponding to the speech frames from the acoustic features. Furthermore, the method may further include determining phoneme data corresponding to the speech frames based at least on the acoustic features. The method may also include determining a quality level of the speech data based on the phoneme data and the feature data, the quality level indicating speech quality of the speech data.

[0005] In a second aspect of the present disclosure, a model generation method is provided. The method may include determining sample acoustic features corresponding to sample speech frames in sample speech data. The method may also include determining feature data and a quality level for each sample phoneme corresponding to the sample speech frame. In addition, the method may include selecting at least two phonemes from the sample speech data as at least a portion of additional sample speech data. Furthermore, the method may include determining an additional quality level for the additional sample speech data based on the quality levels and feature data corresponding to the at least two phonemes. The method may further include training the model based at least on the additional sample speech data and the additional quality level.

[0006] In a third aspect of the present disclosure, a device for processing speech data is provided. The device includes: an acoustic feature determination module configured to determine acoustic features corresponding to speech frames in the speech data; a feature data extraction module configured to extract feature data corresponding to the speech frames from the acoustic features; a phoneme data determination module configured to determine phoneme data corresponding to the speech frames based at least on the acoustic features; and a quality level determination module configured to determine a quality level of the speech data based on the phoneme data and the multi-layer feature data, the quality level indicating speech quality of the speech data.

[0007] In a fourth aspect of the present disclosure, a model generation device is provided. The device includes: a sample acoustic feature determination module configured to determine sample acoustic features corresponding to sample speech frames in sample speech data; a phoneme information determination module configured to determine feature data and a quality level of each sample phoneme corresponding to the sample speech frame; an additional sample speech data determination module configured to select at least two phonemes from the sample speech data as at least a part of additional sample speech data; an additional supervisory information determination module configured to determine an additional quality level of the additional sample speech data based on the quality levels and feature data corresponding to the at least two phonemes; and a model training module configured to train the model based at least on the additional sample speech data and the additional quality level.

[0008] In the fifth aspect of the present disclosure, an electronic device is provided, comprising at least one processor; and a storage device for storing at least one program, wherein when the at least one program is executed by the at least one processor, the at least one processor implements the method according to the first and second aspects of the present disclosure.

[0009] In a sixth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method according to the first and second aspects of the present disclosure is implemented.

[0010] This summary is provided to introduce a selection of concepts in a simplified form that are further described in the detailed description below. This summary is not intended to identify key features or essential features of the disclosure, nor is it intended to limit the scope of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other purposes, features and advantages of the present disclosure will become more apparent by describing the exemplary embodiments of the present disclosure in more detail with reference to the accompanying drawings, wherein the same or similar reference numerals generally represent the same or similar components in the exemplary embodiments of the present disclosure. In the accompanying drawings:

[0012] Figure 1 A schematic diagram illustrating an example environment 100 in which devices and / or methods of embodiments of the present disclosure may be implemented;

[0013] Figure 2 A schematic diagram illustrating a detailed example environment for training and applying a model according to an embodiment of the present disclosure is illustrated;

[0014] Figure 3 illustrates a flow chart of a process 300 for processing speech data according to an embodiment of the present disclosure;

[0015] Figure 4 illustrates a flow chart of a process 400 for determining phoneme data according to an embodiment of the present disclosure;

[0016] Figure 5 A schematic diagram illustrating an example process 500 for generating an acoustic model according to an embodiment of the present disclosure is illustrated;

[0017] Figure 6 illustrates a flow chart of a process 600 for determining a quality level according to an embodiment of the present disclosure;

[0018] Figure 7 A schematic diagram illustrating an example process 700 for determining a quality level of speech data according to an embodiment of the present disclosure is illustrated;

[0019] Figure 8 A flowchart illustrating a process 800 of model generation according to an embodiment of the present disclosure is shown;

[0020] Figure 9 FIG2 illustrates a schematic block diagram of an apparatus 900 for processing voice data according to an embodiment of the present disclosure;

[0021] Figure 10 Illustrated is a schematic block diagram of an example device 1000 suitable for implementing embodiments of the present disclosure.

[0022] In the various drawings, the same or corresponding reference numerals denote the same or corresponding parts. DETAILED DESCRIPTION

[0023] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0024] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0025] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0026] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0027] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0028] The principles of the present disclosure will be described below with reference to several example embodiments shown in the accompanying drawings.

[0029] As used herein, the term "including" and its variations represent open inclusion, i.e., "including but not limited to." Unless otherwise stated, the term "or" means "and / or." The term "based on" means "based at least in part on." The terms "an example embodiment" and "an embodiment" mean "a set of example embodiments." The term "another embodiment" means "a set of additional embodiments." The terms "first," "second," etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0030] As mentioned above, traditional pronunciation quality assessment systems usually use automatic speech recognition technology to extract pronunciation features from the user's (or so-called "second language" user) speech, and use a trained pronunciation quality assessment model to map the extracted pronunciation features into a pronunciation quality score for oral scoring. However, the extraction of pronunciation features and the assessment of pronunciation quality both require a large number of users' second language pronunciation corpora for training. Since users' second language pronunciation is difficult to collect and the annotation cost is high, the scarcity of second language corpora affects the training of the model, thereby limiting the performance of the pronunciation quality assessment system. Therefore, how to build a high-performance pronunciation quality assessment system based on a limited second language pronunciation corpus is a challenge.

[0031] In light of this, embodiments of the present disclosure propose a method for processing speech data. In this method, the acoustic features of each speech frame in the speech data are first determined. Next, at the frame level, feature data for each speech frame is extracted from the acoustic features, and phoneme data for each speech frame is determined based on the acoustic features. Finally, the quality level of the entire speech data is determined based on the determined frame-level feature data and phoneme data. This method implements frame-level speech quality assessment, thereby optimizing the quality level determination process.

[0032] In addition, an embodiment of the present disclosure also proposes a model generation scheme. In this scheme, the sample acoustic features of each sample speech frame in the sample speech data are first determined. Then, at the frame level, the feature data and quality level of each sample phoneme of each sample speech frame are determined. Furthermore, at least two phonemes are selected from the sample speech data to form augmented sample speech data, and based on the quality level and feature data corresponding to the at least two phonemes, the additional quality level of the augmented sample speech data is determined. Finally, the model is trained based on the augmented sample speech data, the additional quality level, and the original training data. The model trained in this way can output more accurate evaluation results.

[0033] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Figure 1 An example environment 100 is shown in which devices and / or methods of embodiments of the present disclosure may be implemented.

[0034] Included in the environment 100 is a computing device 104 for processing speech data 102 to determine a quality level 106 of a user's utterance in the speech data 102 .

[0035] Examples of computing device 104 include, but are not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multi-processor systems, consumer electronics, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, and the like.

[0036] The voice data 102 received by the computing device 104 includes user voice data, and examples of the voice data include, but are not limited to, English voice spoken by a user learning English and Chinese voice spoken by a user learning Chinese. Figure 1 The computing device 104 is shown receiving the voice data 102. This is merely an example and not a limitation of the present disclosure. The computing device 104 may also generate the voice data 102 or the voice data 102 may be stored in a local memory of the computing device 104.

[0037] After obtaining the speech data 102, the computing device 104 can process the speech data 102 to obtain acoustic features 1041 of the speech frames in the speech data. In one example, the acoustic feature is a Mel Frequency Cepstrum Coefficient (MFCC) feature, which is obtained by framing, windowing, and transforming the speech data. In another example, the acoustic feature is any suitable feature that can represent a speech frame, such as a fast Fourier transform feature. The above examples are only used to describe the present disclosure and are not specific limitations of the present disclosure. Those skilled in the art can use any suitable method to obtain the features of the speech data.

[0038] The computing device 104 uses acoustic features 1401 to obtain frame-level feature data 1042 and phoneme data 1043. In this disclosure, a phoneme refers to the smallest unit of speech defined by the natural properties of speech. From an acoustic perspective, a phoneme is the smallest unit of speech defined by sound quality. From a physiological perspective, a phoneme is formed by a single pronunciation action. For example, the English International Phonetic Alphabet (IPA) has 48 phonemes, including 20 vowel phonemes and 28 consonant phonemes.

[0039] For example, for a speech segment, if it includes 50 speech frames, MMFC features for the 50 frames will be generated. Then, the phoneme data corresponding to each of the 50 speech frames is determined. In some embodiments, the use of acoustic features to determine the phoneme data is obtained using a trained acoustic model. In one example, the acoustic model is a DNN-HMM model, and in another example, the acoustic model is a CNN-HMM. The above examples are only used to describe the present disclosure, and are not specific limitations of the present disclosure. Those skilled in the art can set any suitable model to determine the phoneme data.

[0040] The computing device 104 then uses the feature data and corresponding phoneme data of each speech frame to determine a quality level 160 of the speech data 102. In some embodiments, the computing device 104 uses an acoustic model to obtain phonemes and also obtains model-related features from layers prior to the output layer of the acoustic model. The model-related features and phonemes are then used to determine the quality level 160 of the speech data 102. The above examples are merely illustrative of the present disclosure and are not intended to limit the present disclosure.

[0041] The technical solution described above is only for illustration and does not limit the present invention. It should be understood that the system can also be arranged in other ways and connection relationships. In order to explain the principle of the above solution more clearly, the following will refer to Figure 2 Let’s describe the model training and application process in more detail.

[0042] Figure 2FIG2 illustrates a schematic diagram of a detailed example environment 200 for training and applying a model according to an embodiment of the present disclosure. Figure 1 Similarly, the example environment 200 may include a computing device 220, user voice data 210 input to the computing device 220, and a quality level 230 corresponding to the user voice data 210 output from the computing device 220. The difference is that the example environment 200 may include a model training system 260 and a model application system 270. As an example, the model training system 260 and / or the model application system 270 may be implemented in a system such as Figure 1 The computing device 104 shown or Figure 2 The example environment 200 is implemented in the illustrated computing device 220. It should be understood that the structure and functionality of the example environment 200 are described for exemplary purposes only and are not intended to limit the scope of the subject matter described herein. The subject matter described herein can be implemented in different structures and / or functions.

[0043] As previously mentioned, the process of processing the input user voice data 210 to determine the quality level 230 can be divided into two stages: a model training stage and a model application stage. As an example, in the model training stage, the model training system 260 can use the training data set 250 to train the model 240 for performing the corresponding function. It should be understood that the training data set 250 can be a combination of multiple sample data (as input to the model 240) and corresponding annotated supervision information (or referred to as "labels", "true value results"). In the model application stage, the model application system 270 can receive the trained model 240. Thus, the model 240 loaded into the computing device 220 of the model application system 270 can determine the quality level 230 based on the user voice data 210.

[0044] In other embodiments, model 240 can be constructed as a learning network. In some embodiments, the learning network can include multiple networks, each of which can be a multi-layer neural network, which can be composed of a large number of neurons. Through the training process, the corresponding parameters of the neurons in each network can be determined. The parameters of the neurons in these networks are collectively referred to as the parameters of model 240.

[0045] The training process of the model 240 may be performed in an iterative manner until at least some of the parameters of the model 240 converge or until a predetermined number of iterations is reached, thereby obtaining final model parameters.

[0046] The technical solutions described above are only for example and not for limiting the present disclosure. It should be understood that the networks can also be arranged in other ways and connection relationships. Figure 3 Let's describe the voice data processing process in more detail.

[0047] Figure 3 FIGURE 1 illustrates a flow chart of a process 300 for processing speech data according to an embodiment of the present disclosure. In some embodiments, the process 300 may be performed in Figure 1 Implementation in the computing device 104 or other computing devices. Figure 3 Combined with Figure 1 The process 300 of voice data processing according to an embodiment of the present disclosure is described. For ease of understanding, the specific examples mentioned in the following description are exemplary and are not intended to limit the scope of protection of the present disclosure.

[0048] At 302, the computing device 104 may determine acoustic features 1041 corresponding to speech frames in the speech data 102. As an example, the computing device 104 processes the received speech data 102 to determine the acoustic features 1041. In some embodiments, the computing device 104 performs frame processing on the speech data to obtain speech frames. In this way, it is easy to accurately obtain information about the speech data. In some embodiments, the computing device 104 processes the speech data to obtain MFCC features as acoustic features. In some embodiments, the acoustic features are Fast Fourier Transform (FFT) features. The above examples are merely illustrative of the present disclosure and are not intended to be limiting of the present disclosure.

[0049] At 304 , the computing device 104 may extract feature data 1042 corresponding to each speech frame from the acoustic features 1041 . In some embodiments, to extract the feature data, the computing device 104 may extract corresponding feature data from multiple layers of a pre-trained acoustic model as the feature data 1042 .

[0050] As an example, the acoustic model can be a pre-trained wav2vec2 model, which can be obtained through self-learning using a large amount of speech data. The pre-trained wav2vec2 model consists of three parts: an encoder composed of a convolutional neural network, a context processor composed of a Transformer, and a quantizer. Its input is the original speech signal. The encoder can encode a 25ms segment of the speech signal with a sampling rate of 16kHz into a latent vector every 20ms. The context processor can further consider information from other segments in the entire speech on the current segment and further process the latent vector into a context-related segment representation. The quantizer is only used during wav2vec2 pre-training. The pre-trained wav2vec2 model can be provided by another provider or trained by the user using speech data. It should be understood that the multiple layers of the acoustic model mentioned above can refer to multiple layers in the Transformer in the pre-trained wav2vec2 model. In some embodiments, the computing device 104 can extract corresponding feature data from each layer of the acoustic model as feature data 1042.

[0051] At 306, the computing device 104 may determine the phoneme data 1043 corresponding to the speech frame based at least on the acoustic features 1041. In the present disclosure, phoneme refers to the smallest speech unit divided according to the natural properties of speech. As an example, for a segment of speech data, if it includes 50 speech frames, MMFC features for 50 frames will be generated. Then, the phoneme data corresponding to each of the 50 speech frames is determined. Here, it is necessary to determine the phoneme data corresponding to the position of each speech frame, so it is usually necessary to combine the phoneme likelihood value of each speech frame with the text data of the entire speech data to make the determination. Figure 4 The process of determining phoneme data is shown in detail and will be described in detail below.

[0052] At 308, the computing device 104 may determine a quality level 106 for the speech data 102 based on the feature data 1042 and the phoneme data 1043. The quality level 106 generally indicates the speech quality of the speech data 102. In one example, the target quality level is a score. In another example, the target quality level is a different level, such as good, medium, poor, or A, B, C, D. The above examples are merely illustrative of the present disclosure and are not intended to limit the present disclosure. Figure 6 The process of determining the quality level is shown in detail and will be described in detail below.

[0053] Through the above embodiment, the feature data and phoneme data of speech data can be determined at the frame level, and the quality level of the speech data can be determined based on this data. In this way, the robustness of the system and the accuracy of the user's speech evaluation are improved, thereby improving the user experience.

[0054] Further, Figure 4 FIGURE 4 is a flow chart illustrating a process 400 for determining phoneme data according to an embodiment of the present disclosure. In some embodiments, the process 400 may be performed in Figure 1 Implementation in the computing device 104 or other computing devices. Figure 4 Combined with Figure 1 The process 400 for determining phoneme data according to an embodiment of the present disclosure is described. For ease of understanding, the specific examples mentioned in the following description are exemplary and are not intended to limit the scope of protection of the present disclosure.

[0055] At 402, the computing device 104 may apply the acoustic features 1041 to a pre-trained acoustic model. It should be understood that the acoustic model may be trained by the computing device 104, or may be trained by other computing devices and sent to the computing device 104. Furthermore, the computing device 104 may determine the phoneme likelihood values corresponding to each speech frame based on the acoustic features 1041. The phoneme likelihood values are configured to describe the probability that the speech frame corresponds to each phoneme, and may also be used to indicate the phoneme corresponding to the speech frame, for example, corresponding to the phoneme with the maximum probability. In addition to the phoneme likelihood values, the computing device 104 may also obtain text data corresponding to the speech data, that is, a transcribed text of the speech data. Based on the text data and the phoneme likelihood values, the computing device 104 may determine the phoneme data 1043.

[0056] The following combination Figure 5 An example process 500 for generating an acoustic model according to an embodiment of the present disclosure is described. The example process for generating a model may be Figure 1 The process is executed at computing device 104 or other computing device.

[0057] At 502, the computing device 104 can obtain speech data. Then, at 504, the computing device 104 processes the speech data to obtain MFCC features. Next, the computing device 104 uses the MFCC features to train a mixed Gaussian model-hidden Markov model (GMM-HMM) 506. After the GMM-HMM model 506 is trained, the speech data is processed using the GMM-HMM, and then an alignment operation is performed at 508 in combination with the sample text to determine the label of the speech frame. The deep neural network-hidden Markov DNN-HMM model 510 is trained by determining the speech frame label and the corresponding MFCC features. The trained DNN-HMM model can determine the corresponding phoneme based on the MFCC features. Therefore, the trained DNN-HMM model can be used as an acoustic model.

[0058] like Figure 5 As shown, the DNN-HMM includes multiple layers 512, 514, 516 for DNN and a layer 518 for HMM, where layer 512 is the input layer, layer 516 is the output layer, and layer 514 is the hidden layer, which can be one or more layers. Figure 5 This is merely an example and is not intended to limit the present disclosure. An acoustic model such as a CNN-HMM or any other suitable acoustic model may also be trained.

[0059] Combined with the above Figure 5 An example process 500 for generating an acoustic model according to an embodiment of the present disclosure is described. Figure 6 A schematic diagram depicting a process 600 for determining a quality level according to an embodiment of the present disclosure is provided. The process 600 may be performed in Figure 1 The process is executed at computing device 104 or other computing device.

[0060] At 602, computing device 104 may apply the phoneme data and feature data to a pre-trained quality rating determination model. At 604, computing device 104 may weight the feature data of the multiple layers based on a predetermined attention mechanism. It should be understood that the weights under the attention mechanism will generally change with each iteration of training. Alternatively or additionally, the feature data of each layer may be weighted using predetermined weights. At 606, computing device 104 may determine a quality rating based on the weighted feature data and phoneme data.

[0061] The following combination Figure 7 A diagram depicts an example process 700 for determining a quality level of speech data according to an embodiment of the present disclosure.

[0062] The input audio 702 is fed into a feature extraction module 704 for processing. The feature extraction module 704 obtains acoustic features from the speech frame of the input audio. The acoustic features are then input into an acoustic model 706 to obtain phoneme likelihood values 708 and features 710 extracted from each layer in the acoustic model 706. The acoustic model can be a pre-trained wav2vec2 model. Therefore, the wav2vec2 model can be trained to determine the phoneme likelihood value of each phoneme and the features of each layer in the wav2vec2 model from the acoustic features. The phoneme likelihood value 708 describes the probability that the speech frame corresponds to each phoneme, which can also be used to indicate the phoneme to which the speech frame corresponds, for example, the phoneme with the maximum probability.

[0063] Then, the phoneme likelihood value 708 and the input text 712 associated with the input audio are input to a decoder 714. The decoder 714 matches the multiple phoneme likelihood values of the multiple speech frames of the input audio with the phoneme sequence of the input text to determine one or more speech frames corresponding to each phoneme in the phoneme sequence in the text, thereby forming a phoneme timestamp 716. The phoneme timestamp 716 includes a phoneme identifier and the corresponding one or more speech frames. Furthermore, the computing device 104 can input each layer of features 710 and the phoneme timestamp 716 into a quality level determination model 718 to determine a quality level 720. In this way, the acoustic information and phonemes of the speech data can be effectively utilized, the accuracy of the evaluation and the system performance can be improved, and the user experience can be improved.

[0064] In some embodiments, computing device 104 receives trained acoustic models from other computing devices.

[0065] In some embodiments, the computing device 104 trains the acoustic model 706. As an example, the computing device 104 can obtain a first set of sample speech data and pre-train the acoustic model using the first set of sample speech data. As an example, the first set of sample speech data can be approximately 60,000 hours of native language unlabeled pronunciation data. As an example, the first set of sample speech data can be native language unlabeled pronunciation data with a duration greater than 60,000 hours. As an example, the first set of sample speech data can be approximately 10,000 to 100,000 hours of native language unlabeled pronunciation data. As an example, the first set of sample speech data can be approximately 1,000 to 500,000 hours of native language unlabeled pronunciation data.

[0066] Furthermore, the computing device 104 may obtain a second set of sample speech data and corresponding sample text, and fine-tune the acoustic model using the second set of sample speech data and corresponding sample text. As an example, the second set of sample speech data may be approximately 1,000 hours of native language audio data and 10 hours of second language audio data (i.e., the user's audio data) and the corresponding transcribed texts (i.e., the text corresponding to these audio data). As an example, the second set of sample speech data may be approximately 500 to 5,000 hours of native language audio data and 5 to 50 hours of second language audio data and the corresponding transcribed texts.

[0067] In some embodiments, the audio Mel-frequency cepstral coefficient features are first extracted, the Gaussian mixture model-hidden Markov model is trained, and the audio is aligned using the transcribed text to obtain frame-level labels. The wav2vec2 model is trained as an acoustic model using the frame-level labels and the original audio signal. It should be understood that the above description of the duration of the audio data is only exemplary, and these data can be adjusted according to actual conditions and needs. In this way, only a small amount (for example, the above 10 hours) of second language audio data needs to be expertly labeled, thereby saving the labeling cost and improving the training quality of the model.

[0068] In some embodiments, the computing device 104 trains the quality level determination model 718. As an example, the computing device 104 can use a third set of sample voice data and corresponding follow-up text to pre-train the quality level determination model. As an example, the third set of sample voice data can be the above-mentioned 10 hours of second language audio data or other second language audio data and the follow-up text corresponding to these data (i.e., the recognition text of the user's follow-up voice). Specifically, the acoustic model and decoder can be used to perform feature mapping and alignment on the limited above-mentioned 10 hours of scored training data to obtain the output features of each layer of the intermediate layer in the acoustic model for each audio, the phoneme likelihood matrix of the last layer, and the phoneme timestamp. For the 10 hours of real data, the GOP (goodness of pronunciation) score of each phoneme in the sentence is calculated using the phoneme likelihood matrix and the phoneme timestamp, and then the sentence-level GOP score is obtained by averaging. Using this score as supervision information and the output of each layer of the intermediate layer of the acoustic model as a feature, the feature-weighted quality level determination model 718 is trained.

[0069] The computing device 104 may then obtain a fourth set of sample speech data, the corresponding follow-up text, and the sample quality ratings annotated by the expert, and use the fourth set of sample speech data, the corresponding follow-up text, and the sample quality ratings annotated by the expert to fine-tune the quality rating determination model. The fourth set of sample speech data may include approximately four hours or other lengths of second language audio, its follow-up text, and the scores assigned by the language expert to each audio according to the pronunciation assessment rules.

[0070] In some embodiments, training data augmentation may be performed through the following process. For example, the computing device 104 may determine sample acoustic features corresponding to the sample speech frames in the third set of sample speech data, and determine multi-layer feature data and quality levels for each sample phoneme corresponding to the sample speech frames, thereby forming a feature pool consisting of a plurality of phonemes, their corresponding quality levels (e.g., GOP scores), and their corresponding multi-layer feature data.

[0071] Furthermore, the computing device 104 can arbitrarily select at least two phonemes from the third set of sample speech data to form at least a portion of the augmented sample speech data. Thus, the computing device 104 can determine the additional quality level of the augmented sample speech data based on the quality levels corresponding to the at least two phonemes and the multi-layer feature data. As an example, after selecting at least two phonemes from the above-mentioned feature pool, the feature data corresponding to these phonemes can be combined to form augmented feature data, and the augmented quality level can be determined based on the quality levels corresponding to these phonemes. Thus, the computing device 104 can train a quality level determination model based at least on the augmented feature data corresponding to the augmented sample speech data and the determined augmented quality level.

[0072] In certain embodiments, in order to determine the quality level of each sample phoneme corresponding to the sample speech frame, the computing device 104 may apply the sample acoustic features to a pre-trained acoustic model to determine a sample phoneme likelihood value corresponding to the sample speech frame and multi-layer feature data for each sample phoneme, and may determine a sample phoneme timestamp based on the follow-up text of the third set of sample speech data and the sample phoneme likelihood value. Furthermore, the computing device 104 may determine the quality level of each sample phoneme in the third set of sample speech data based on the sample phoneme likelihood value and the sample phoneme timestamp, and select at least two phonemes from the third set of sample speech data as at least a portion of the additional sample speech data. Further, the computing device 104 may determine the quality level of the additional sample speech data based on the quality levels corresponding to the at least two phonemes and the multi-layer feature data.

[0073] Figure 8 FIGURE 8 is a flow chart illustrating a process 800 of model generation according to an embodiment of the present disclosure. Figure 8 A schematic diagram illustrating a process 800 for model generation according to an embodiment of the present disclosure is provided. The process 800 may be performed in Figure 1 The process is executed at computing device 104 or other computing device.

[0074] At 802, the computing device 104 may determine a sample acoustic feature corresponding to a sample speech frame in the sample speech data. This process is similar to the model application process and will not be described in detail here.

[0075] At 804, the computing device 104 may determine feature data and a quality level for each sample phoneme corresponding to the sample speech frame. In some embodiments, the computing device 104 may form a feature pool consisting of a plurality of phonemes, their corresponding quality levels (e.g., GOP scores), and their corresponding multi-layer feature data by determining the feature data and quality level for each sample phoneme in the training data.

[0076] At 806, the computing device 104 may select at least two phonemes from the sample speech data as at least a portion of the additional sample speech data. For example, the computing device 104 may arbitrarily select at least two phonemes from the third set of sample speech data to form at least a portion of the augmented sample speech data. For example, after selecting at least two phonemes from the feature pool, the feature data corresponding to these phonemes may be combined to form augmented feature data, and an augmented quality level may be determined based on the quality levels corresponding to these phonemes. Thus, the computing device 104 may train a quality level determination model based at least on the augmented feature data corresponding to the augmented sample speech data and the determined augmented quality level.

[0077] At 808, computing device 104 may determine an additional quality level for the additional sample speech data based on the quality levels and feature data corresponding to the at least two phonemes. Furthermore, at 810, computing device 104 may train the model based at least on the additional sample speech data and the additional quality level. It should be understood that the data augmentation may be performed on the sample speech data itself or on its corresponding features. In other words, computing device 104 may perform model training based on the feature data and quality level corresponding to the additional sample speech data, or based on the additional sample speech data itself and the quality level as supervisory information.

[0078] In certain embodiments, to determine the quality level of each sample phoneme corresponding to a sample speech frame, the computing device 104 may apply the sample acoustic features to a pre-trained acoustic model to determine a sample phoneme likelihood value corresponding to the sample speech frame and multi-layer feature data for each sample phoneme. Furthermore, the computing device 104 may determine a sample phoneme timestamp based on the follow-up text of the sample speech data and the sample phoneme likelihood value. The computing device 104 may then determine the quality level of each sample phoneme in the sample speech data based on the sample phoneme likelihood value and the sample phoneme timestamp, thereby modifying the feature pool.

[0079] Furthermore, the computing device 104 may select at least two phonemes from the sample speech data in the feature pool as at least a part of the additional sample speech data, and determine the quality level of the additional sample speech data based on the quality levels corresponding to the at least two phonemes and the multi-layer feature data.

[0080] Through the above-mentioned embodiments, the present disclosure first realizes the determination of feature data and phoneme data at the frame level, thereby making the determined quality level more accurate. Furthermore, the present disclosure utilizes open-source large-scale unlabeled native language data to pre-train the wav2vec2 model, and then uses a small amount of labeled native language and second language data for fine-tuning, thereby further improving the robustness of the system. In addition, the present disclosure effectively utilizes the features output by each layer in the acoustic model, thereby making the representation of the extracted user's second language speech more comprehensive. More importantly, the present disclosure creates a phoneme-level feature library, so that more training data can be augmented in a free combination manner, and the features and quality levels of these augmented training data can be determined, thereby expanding the training data set at a low cost and improving the quality of the model.

[0081] The present disclosure also provides a device for processing voice data. Specifically, Figure 9 FIG. 1 shows a schematic diagram of an apparatus 900 for processing voice data according to an embodiment of the present disclosure. Figure 9 As shown, the apparatus 900 may include at least: an acoustic feature determination module 902, configured to determine acoustic features corresponding to speech frames in the speech data; a feature data extraction module 904, configured to extract feature data corresponding to the speech frames from the acoustic features; a phoneme data determination module 906, configured to determine phoneme data corresponding to the speech frames based at least on the acoustic features; and a quality level determination module 908, configured to determine the quality level of the speech data based on the phoneme data and the multi-layer feature data, wherein the quality level indicates the speech quality of the speech data.

[0082] In some embodiments, the phoneme data determination module 906 can be configured to: apply the acoustic features to a pre-trained acoustic model; determine the phoneme likelihood value corresponding to the speech frame based on the acoustic features; and determine the phoneme data based on the text data corresponding to the speech data and the phoneme likelihood value.

[0083] In some embodiments, the feature data extraction module 904 may be configured to extract corresponding feature data from multiple layers of a pre-trained acoustic model as the feature data.

[0084] In some embodiments, the quality level determination module 908 can be configured to: apply the phoneme data and the feature data to a pre-trained quality level determination model; perform weighted processing on the feature data of multiple layers based on a predetermined attention mechanism; and determine the quality level based on the weighted feature data and the phoneme data.

[0085] In some embodiments, the device 900 may also include: a first sample data acquisition module, configured to acquire a first set of sample speech data; a first pre-training module, configured to pre-train the acoustic model using the first set of sample speech data; a second sample data acquisition module, configured to acquire a second set of sample speech data and corresponding sample text; and a first fine-tuning module, configured to fine-tune the acoustic model using the second set of sample speech data and the sample text.

[0086] In some embodiments, the second set of sample voice data includes native language voice data having a first predetermined duration and language learning voice data having a second predetermined duration.

[0087] In some embodiments, the device 900 may also include: a third sample data acquisition module, configured to acquire a third group of sample voice data and corresponding follow-up text; a second pre-training module, configured to pre-train the quality level determination model using the third group of sample voice data and the corresponding follow-up text; a fourth sample data acquisition module, configured to acquire a fourth group of sample voice data, corresponding follow-up text and sample quality levels annotated by experts; and a second fine-tuning module, configured to fine-tune the quality level determination model using the fourth group of sample voice data, corresponding follow-up text and the sample quality levels annotated by experts.

[0088] In some embodiments, the third set of sample voice data is language learning voice data having a third predetermined duration, and the fourth set of sample voice data is language learning voice data having a fourth predetermined duration.

[0089] In some embodiments, the device 900 may also include: a sample acoustic feature determination module, configured to determine the sample acoustic features corresponding to the sample speech frames in the third group of sample speech data; a phoneme information determination module, configured to determine the multi-layer feature data and quality level of each sample phoneme corresponding to the sample speech frame; an additional sample speech data determination module, configured to select at least two phonemes from the third group of sample speech data as at least a part of the additional sample speech data; an additional supervisory information determination module, configured to determine the additional quality level of the additional sample speech data based on the quality levels and multi-layer feature data corresponding to the at least two phonemes; and a model training module, configured to train the quality level determination model based at least on the additional sample speech data and the additional quality level.

[0090] In some embodiments, the phoneme information determination module 906 is configured to: apply the sample acoustic features to a pre-trained acoustic model to determine a sample phoneme likelihood value corresponding to the sample speech frame and multi-layer feature data for each sample phoneme; determine a sample phoneme timestamp based on the follow-up text of the third group of sample speech data and the sample phoneme likelihood value; determine a quality level of each sample phoneme in the third group of sample speech data based on the sample phoneme likelihood value and the sample phoneme timestamp; select at least two phonemes from the third group of sample speech data as at least a part of the additional sample speech data; and determine the quality level of the additional sample speech data based on the quality levels corresponding to the at least two phonemes and the multi-layer feature data.

[0091] In addition, although not shown, the present disclosure also provides a model generation device, including: a sample acoustic feature determination module, configured to determine the sample acoustic features corresponding to the sample speech frame in the sample speech data; a phoneme information determination module, configured to determine the feature data and quality level of each sample phoneme corresponding to the sample speech frame; an additional sample speech data determination module, configured to select at least two phonemes from the sample speech data as at least a part of the additional sample speech data; an additional supervisory information determination module, configured to determine the additional quality level of the additional sample speech data based on the quality level and feature data corresponding to the at least two phonemes; and a model training module, configured to train the model based at least on the additional sample speech data and the additional quality level.

[0092] In some embodiments, the phoneme information determination module is configured to include: applying the sample acoustic features to a pre-trained acoustic model to determine a sample phoneme likelihood value corresponding to the sample speech frame and multi-layer feature data for each sample phoneme; determining a sample phoneme timestamp based on the follow-up text of the sample speech data and the sample phoneme likelihood value; determining a quality level of each sample phoneme in the sample speech data based on the sample phoneme likelihood value and the sample phoneme timestamp; selecting at least two phonemes from the sample speech data as at least a part of the additional sample speech data; and determining the quality level of the additional sample speech data based on the quality levels corresponding to the at least two phonemes and the multi-layer feature data.

[0093] Figure 10 A block diagram of a computing device 1000 capable of implementing various embodiments of the present disclosure is shown. Electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit implementations of the present disclosure described and / or claimed herein.

[0094] like Figure 10 As shown, the device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0095] Various components in device 1000 are connected to I / O interface 1005, including an input unit 1006, such as a keyboard, mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, optical disk, etc.; and a communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0096] The computing unit 1001 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1001 performs the various methods and processes described above, such as processes 300, 500, 600, and 800. For example, in some embodiments, the processes 300, 500, 600, and 800 may be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into RAM 1003 and executed by computing unit 1001, one or more steps of processes 300, 500, 600, and 800 described above may be performed. Alternatively, in other embodiments, computing unit 1001 may be configured to perform processes 300, 500, 600, and 800 in any other appropriate manner (e.g., by means of firmware).

[0097] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0098] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0099] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0100] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0101] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0102] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

[0103] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0104] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

[0105] According to one or more embodiments of the present disclosure, Example 1. A method for processing speech data, comprising: determining acoustic features corresponding to speech frames in the speech data; extracting feature data corresponding to the speech frames from the acoustic features; determining phoneme data corresponding to the speech frames based at least on the acoustic features; and determining a quality level of the speech data based on the phoneme data and the feature data, the quality level indicating the speech quality of the speech data.

[0106] Example 2. A method according to Example 1, wherein determining the phoneme data includes: applying the acoustic features to a pre-trained acoustic model; determining a phoneme likelihood value corresponding to the speech frame based on the acoustic features; and determining the phoneme data based on text data corresponding to the speech data and the phoneme likelihood value.

[0107] Example 3. The method according to Example 1, wherein extracting the feature data comprises: extracting corresponding feature data from multiple layers of a pre-trained acoustic model as the feature data.

[0108] Example 4. A method according to Example 3, wherein determining the quality level includes: applying the phoneme data and the feature data to a pre-trained quality level determination model; performing weighted processing on the feature data of multiple layers based on a predetermined attention mechanism; and determining the quality level based on the weighted feature data and the phoneme data.

[0109] Example 5. The method according to Example 2 further includes: obtaining a first set of sample speech data; pre-training the acoustic model using the first set of sample speech data; obtaining a second set of sample speech data and corresponding sample text; and fine-tuning the acoustic model using the second set of sample speech data and the sample text.

[0110] Example 6. The method of Example 5, wherein the second set of sample speech data includes native language speech data having a first predetermined duration and language learning speech data having a second predetermined duration.

[0111] Example 7. The method according to Example 4 further includes: obtaining a third group of sample voice data and corresponding follow-up text; pre-training the quality level determination model using the third group of sample voice data and the corresponding follow-up text; obtaining a fourth group of sample voice data, corresponding follow-up text and sample quality levels annotated by experts; and fine-tuning the quality level determination model using the fourth group of sample voice data, corresponding follow-up text and the sample quality levels annotated by experts.

[0112] Example 8. The method of Example 7, wherein the third set of sample voice data is language learning voice data having a third predetermined duration, and the fourth set of sample voice data is language learning voice data having a fourth predetermined duration.

[0113] Example 9. The method according to Example 8 further includes: determining sample acoustic features corresponding to sample speech frames in the third group of sample speech data; determining multi-layer feature data and quality levels for each sample phoneme corresponding to the sample speech frame; selecting at least two phonemes from the third group of sample speech data as at least a part of additional sample speech data; determining additional quality levels of the additional sample speech data based on the quality levels and multi-layer feature data corresponding to the at least two phonemes; and training the quality level determination model based at least on the additional sample speech data and the additional quality levels.

[0114] Example 10. A method according to Example 9, wherein determining the quality level of each sample phoneme corresponding to the sample speech frame includes: applying the sample acoustic features to a pre-trained acoustic model to determine a sample phoneme likelihood value corresponding to the sample speech frame and multi-layer feature data for each sample phoneme; determining a sample phoneme timestamp based on the follow-up text of the third group of sample speech data and the sample phoneme likelihood value; determining the quality level of each sample phoneme in the third group of sample speech data based on the sample phoneme likelihood value and the sample phoneme timestamp; selecting at least two phonemes from the third group of sample speech data as at least a part of the additional sample speech data; and determining the quality level of the additional sample speech data based on the quality levels corresponding to the at least two phonemes and the multi-layer feature data.

[0115] Example 11. A model generation method, comprising: determining sample acoustic features corresponding to sample speech frames in sample speech data; determining feature data and a quality level for each sample phoneme corresponding to the sample speech frame; selecting at least two phonemes from the sample speech data as at least a part of additional sample speech data; determining an additional quality level for the additional sample speech data based on the quality levels and feature data corresponding to the at least two phonemes; and training the model based at least on the additional sample speech data and the additional quality level.

[0116] 12. A method according to Example 11, wherein determining the quality level of each sample phoneme corresponding to the sample speech frame includes: applying the sample acoustic features to a pre-trained acoustic model to determine a sample phoneme likelihood value corresponding to the sample speech frame and multi-layer feature data for each sample phoneme; determining a sample phoneme timestamp based on the follow-up text of the sample speech data and the sample phoneme likelihood value; determining the quality level of each sample phoneme in the sample speech data based on the sample phoneme likelihood value and the sample phoneme timestamp; selecting at least two phonemes from the sample speech data as at least a part of the additional sample speech data; and determining the quality level of the additional sample speech data based on the quality levels corresponding to the at least two phonemes and the multi-layer feature data.

[0117] 13. A device for processing speech data, comprising: an acoustic feature determination module, configured to determine acoustic features corresponding to speech frames in the speech data; a feature data extraction module, configured to extract feature data corresponding to the speech frames from the acoustic features; a phoneme data determination module, configured to determine phoneme data corresponding to the speech frames based at least on the acoustic features; and a quality level determination module, configured to determine the quality level of the speech data based on the phoneme data and the multi-layer feature data, the quality level indicating the speech quality of the speech data.

[0118] Example 14. An apparatus according to Example 13, wherein the phoneme data determination module is configured to: apply the acoustic features to a pre-trained acoustic model; determine a phoneme likelihood value corresponding to the speech frame based on the acoustic features; and determine the phoneme data based on text data corresponding to the speech data and the phoneme likelihood value.

[0119] Example 15. The apparatus according to Example 13, wherein the feature data extraction module is configured to extract corresponding feature data from multiple layers of a pre-trained acoustic model as the feature data.

[0120] Example 16. An apparatus according to Example 15, wherein the quality level determination module is configured to: apply the phoneme data and the feature data to a pre-trained quality level determination model; perform weighted processing on the feature data of multiple layers based on a predetermined attention mechanism; and determine the quality level based on the weighted feature data and the phoneme data.

[0121] Example 17. The apparatus according to Example 14 further includes: a first sample data acquisition module, configured to acquire a first set of sample speech data; a first pre-training module, configured to pre-train the acoustic model using the first set of sample speech data; a second sample data acquisition module, configured to acquire a second set of sample speech data and corresponding sample text; and a first fine-tuning module, configured to fine-tune the acoustic model using the second set of sample speech data and the sample text.

[0122] Example 18. The apparatus of Example 17, wherein the second set of sample speech data comprises native language speech data having a first predetermined duration and language learning speech data having a second predetermined duration.

[0123] Example 19. The device according to Example 16 further includes: a third sample data acquisition module, configured to acquire a third group of sample voice data and corresponding follow-up text; a second pre-training module, configured to pre-train the quality level determination model using the third group of sample voice data and the corresponding follow-up text; a fourth sample data acquisition module, configured to acquire a fourth group of sample voice data, corresponding follow-up text and sample quality levels annotated by experts; and a second fine-tuning module, configured to fine-tune the quality level determination model using the fourth group of sample voice data, corresponding follow-up text and the sample quality levels annotated by experts.

[0124] Example 20. The apparatus of Example 19, wherein the third set of sample voice data is language learning voice data having a third predetermined duration, and the fourth set of sample voice data is language learning voice data having a fourth predetermined duration.

[0125] Example 21. The apparatus according to Example 20 further includes: a sample acoustic feature determination module configured to determine sample acoustic features corresponding to sample speech frames in the third group of sample speech data; a phoneme information determination module configured to determine multi-layer feature data and quality levels of each sample phoneme corresponding to the sample speech frame; an additional sample speech data determination module configured to select at least two phonemes from the third group of sample speech data as at least a part of the additional sample speech data; an additional supervisory information determination module configured to determine additional quality levels of the additional sample speech data based on the quality levels and multi-layer feature data corresponding to the at least two phonemes; and a model training module configured to train the quality level determination model based at least on the additional sample speech data and the additional quality levels.

[0126] Example 22. An apparatus according to Example 21, wherein the phoneme information determination module is configured to: apply the sample acoustic features to a pre-trained acoustic model to determine a sample phoneme likelihood value corresponding to the sample speech frame and multi-layer feature data for each sample phoneme; determine a sample phoneme timestamp based on the follow-up text of the third group of sample speech data and the sample phoneme likelihood value; determine a quality level of each sample phoneme in the third group of sample speech data based on the sample phoneme likelihood value and the sample phoneme timestamp; select at least two phonemes from the third group of sample speech data as at least a part of the additional sample speech data; and determine the quality level of the additional sample speech data based on the quality levels corresponding to the at least two phonemes and the multi-layer feature data.

[0127] 23. A model generation device comprising:

[0128] a sample acoustic feature determination module configured to determine a sample acoustic feature corresponding to a sample speech frame in the sample speech data;

[0129] a phoneme information determination module configured to determine feature data and a quality level of each sample phoneme corresponding to the sample speech frame;

[0130] an additional sample voice data determining module, configured to select at least two phonemes from the sample voice data as at least a part of the additional sample voice data;

[0131] an additional supervisory information determination module configured to determine an additional quality level of the additional sample speech data based on the quality levels and feature data corresponding to the at least two phonemes; and

[0132] The model training module is configured to train the model based on at least the additional sample speech data and the additional quality level.

[0133] 24. An apparatus according to Example 23, wherein the phoneme information determination module is configured to include: applying the sample acoustic features to a pre-trained acoustic model to determine a sample phoneme likelihood value corresponding to the sample speech frame and multi-layer feature data for each sample phoneme; determining a sample phoneme timestamp based on the follow-up text of the sample speech data and the sample phoneme likelihood value; determining a quality level of each sample phoneme in the sample speech data based on the sample phoneme likelihood value and the sample phoneme timestamp; selecting at least two phonemes from the sample speech data as at least a part of the additional sample speech data; and determining the quality level of the additional sample speech data based on the quality levels corresponding to the at least two phonemes and the multi-layer feature data.

[0134] 25. An electronic device comprising:

[0135] at least one processor; and

[0136] A storage device for storing at least one program, when the at least one program is executed by the at least one processor, so that the at least one processor implements the method according to any one of Examples 1-12.

[0137] 26. A computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the method according to any one of Examples 1-12.

Claims

1. A method for processing speech data, comprising: determining acoustic features corresponding to speech frames in the speech data; Extracting feature data corresponding to the speech frame from the acoustic features; determining phoneme data corresponding to the speech frame based at least on the acoustic feature; as well as determining a quality level of the speech data based on the phoneme data and the feature data, the quality level indicating speech quality of the speech data, The extracting of the characteristic data comprises: Extracting corresponding feature data from multiple layers of a pre-trained acoustic model as the feature data, Wherein determining the quality level comprises: Applying the phoneme data and the feature data to a pre-trained quality level determination model; Weighting the multiple layers of feature data based on a predetermined attention mechanism; and The quality level is determined based on the weighted feature data and the phoneme data.

2. The method of claim 1 , wherein determining the phoneme data comprises: Applying the acoustic features to a pre-trained acoustic model; Determining a phoneme likelihood value corresponding to the speech frame based on the acoustic feature; as well as The phoneme data is determined based on text data corresponding to the speech data and the phoneme likelihood value.

3. The method according to claim 2, further comprising: Obtaining a first set of sample voice data; Pre-training the acoustic model using the first set of sample speech data; Obtaining a second set of sample speech data and corresponding sample text; as well as The acoustic model is fine-tuned using the second set of sample speech data and the sample text. 4 . The method according to claim 3 , wherein the second set of sample speech data comprises native language speech data having a first predetermined duration and language learning speech data having a second predetermined duration.

5. The method according to claim 1, further comprising: Obtain the third set of sample voice data and the corresponding follow-up text; Pre-training the quality level determination model using the third set of sample speech data and corresponding follow-up text; Obtain the fourth set of sample speech data, the corresponding follow-up text, and the sample quality level annotated by the expert; as well as The quality grade determination model is fine-tuned using the fourth set of sample speech data, the corresponding follow-up text, and the sample quality grades annotated by experts. 6 . The method according to claim 5 , wherein the third set of sample voice data is language learning voice data having a third predetermined duration, and the fourth set of sample voice data is language learning voice data having a fourth predetermined duration.

7. The method according to claim 6, further comprising: determining sample acoustic features corresponding to sample speech frames in the third set of sample speech data; determining multi-layer feature data and a quality level for each sample phoneme corresponding to the sample speech frame; selecting at least two phonemes from the third set of sample speech data as at least a portion of the additional sample speech data; determining an additional quality level for the additional sample speech data based on the quality levels corresponding to the at least two phonemes and the multi-layer feature data; as well as The quality level determination model is trained based on at least the additional sample speech data and the additional quality levels.

8. The method of claim 7, wherein determining the quality level of each sample phoneme corresponding to the sample speech frame comprises: Applying the sample acoustic features to a pre-trained acoustic model to determine a sample phoneme likelihood value corresponding to the sample speech frame and multi-layer feature data for each sample phoneme; Determining a sample phoneme timestamp based on the follow-up text of the third set of sample voice data and the sample phoneme likelihood value; determining a quality level of each sample phoneme in the third set of sample speech data based on the sample phoneme likelihood value and the sample phoneme timestamp; selecting at least two phonemes from the third set of sample speech data as at least a portion of the additional sample speech data; and The quality level of the additional sample speech data is determined based on the quality levels corresponding to the at least two phonemes and the multi-layer feature data.

9. A device for processing voice data, comprising: an acoustic feature determination module, configured to determine acoustic features corresponding to speech frames in the speech data; a feature data extraction module, configured to extract feature data corresponding to the speech frame from the acoustic features; a phoneme data determination module, configured to determine phoneme data corresponding to the speech frame based at least on the acoustic feature; as well as a quality level determination module configured to determine a quality level of the speech data based on the phoneme data and the feature data, wherein the quality level indicates the speech quality of the speech data; The feature data extraction module is further configured to: Extracting corresponding feature data from multiple layers of a pre-trained acoustic model as the feature data, The quality level determination module is further configured to: Applying the phoneme data and the feature data to a pre-trained quality level determination model; Weighting the multiple layers of feature data based on a predetermined attention mechanism; as well as The quality level is determined based on the weighted feature data and the phoneme data.

10. An electronic device comprising: at least one processor; as well as A storage device for storing at least one program, wherein when the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Speech scoring method and device, electronic device, and storage medium

    CN109256152A

  • Pronunciation defect recognition model training method and pronunciation defect recognition method

    CN112687291A

  • Voice evaluation method and device, equipment and storage medium

    CN114627896A