Voice interaction method and device and storage medium

By combining pronunciation review and speech synthesis technology, the feedback voice of the specified tone is generated, which solves the problem of low voice selectivity for feedback voice in existing pronunciation review products, and realizes personalized customization and high selectivity.

CN119943026APending Publication Date: 2025-05-06GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202311452253.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing pronunciation review products have low feedback voice selection and fixed or limited tone.

Method used

By combining pronunciation evaluation technology and speech synthesis technology, the server scores the first audio data and generates a specified tone voice corresponding to the score. The speech synthesis model is trained using user-specified training data to realize personalized customization of feedback speech.

Benefits of technology

It improves the selectability of feedback voice, enables users to select tones according to their personal preferences, and enhances the fun and user experience of the pronunciation review product.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943026A_ABST
    Figure CN119943026A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice interaction method and device, equipment and a storage medium, and belongs to the technical field of pronunciation evaluation. The method comprises the following steps: acquiring a pronunciation score of first audio data and a voice feedback text corresponding to the pronunciation score; calling a voice synthesis model to synthesize a specified tone voice corresponding to the voice feedback text; wherein the speech synthesis model is obtained by training second audio data of a specified speaker; and feeding back the specified timbre voice to the user side, so that the user side plays the specified timbre voice. According to the method and the device, personalized customization of the feedback voice is realized, and the interestingness of a pronunciation evaluation product is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of pronunciation evaluation, and in particular to a voice interaction method, device and storage medium. Background Art

[0002] With the continuous development of science and technology, voice-based intelligent interaction has gradually spread to various fields. Pronunciation evaluation is a technology that automatically scores users' spoken expressions, which can be used to determine whether the user's spoken expression is standard. For example, the application of pronunciation evaluation technology in the field of education can help users improve their oral expression skills.

[0003] At present, there are many products with pronunciation evaluation functions on the market, such as a certain brand of learning machine, which is mainly aimed at evaluating children's English and Chinese pronunciation. Usually, this type of pronunciation evaluation product can first score the user's pronunciation, and after obtaining the pronunciation evaluation score, it can feedback the specific score to the user through voice, and can also give corresponding praise or encouragement.

[0004] However, the existing pronunciation evaluation products only use one or more fixed timbres for the feedback voice, which results in low selectivity. Summary of the invention

[0005] The embodiments of the present application provide a voice interaction method, device and storage medium, which can solve the problem of low selectivity of feedback voice in existing pronunciation evaluation products. To solve the above problems, the technical solutions provided by the embodiments of the present application are as follows:

[0006] In a first aspect, an embodiment of the present application provides a voice interaction method, the method comprising:

[0007] Calling a pronunciation evaluation model to score the first audio data to obtain a pronunciation score of the user, and determining a voice feedback text corresponding to the pronunciation score;

[0008] Calling a speech synthesis model to synthesize a speech with a specified timbre corresponding to the speech feedback text; wherein the speech synthesis model is trained using second audio data of a specified speech;

[0009] Feedback the specified timbre voice to the user terminal so that the user terminal plays the specified timbre voice.

[0010] In a second aspect, an embodiment of the present application provides a voice interaction server, comprising:

[0011] A pronunciation evaluation module, configured to call a pronunciation evaluation model to score the first audio data, obtain a pronunciation score of the user, and determine a voice feedback text corresponding to the pronunciation score;

[0012] A speech synthesis module, used for calling a speech synthesis model to synthesize a speech of a specified timbre corresponding to the speech feedback text; wherein the speech synthesis model is trained using the second audio data of the specified speaker;

[0013] The designated timbre feedback module is used to feed back the designated timbre voice to the user terminal so that the user terminal plays the designated timbre voice.

[0014] In a third aspect, an embodiment of the present application provides a voice interaction method, the method comprising:

[0015] Acquire first audio data, and upload the first audio data to a server;

[0016] Receive and play the specified tone voice fed back by the server;

[0017] The specified timbre voice is generated by the server using the following steps:

[0018] Calling a pronunciation evaluation model to score the first audio data to obtain a pronunciation score of the user, and determining a voice feedback text corresponding to the pronunciation score;

[0019] A speech synthesis model is called to synthesize a speech with a specified timbre corresponding to the speech feedback text; wherein the speech synthesis model is trained using the second audio data of a specified speaker.

[0020] In a fourth aspect, an embodiment of the present application provides a voice interaction user terminal, including:

[0021] An audio acquisition module, used to acquire first audio data and upload the first audio data to a server;

[0022] A designated timbre receiving module, used for receiving and playing the designated timbre voice fed back by the server;

[0023] The specified timbre voice is generated by the server using the following steps:

[0024] Calling a pronunciation evaluation model to score the first audio data to obtain a pronunciation score of the user, and determining a voice feedback text corresponding to the pronunciation score;

[0025] Calling a speech synthesis model to synthesize a speech with a specified timbre corresponding to the speech feedback text; wherein,

[0026] The speech synthesis model is trained using second audio data of a designated speaker.

[0027] In a fifth aspect, an embodiment of the present application provides a voice interaction device, the voice interaction device comprising:

[0028] An audio acquisition module, used to acquire first audio data;

[0029] A pronunciation evaluation module, configured to call a pronunciation evaluation model to score the first audio data, obtain a pronunciation score of the user, and determine a voice feedback text corresponding to the pronunciation score;

[0030] A speech synthesis module, used for calling a speech synthesis model to synthesize a speech of a specified timbre corresponding to the speech feedback text; wherein the speech synthesis model is trained using the second audio data of the specified speaker;

[0031] The voice playing module is used to play the voice with the specified timbre.

[0032] In a sixth aspect, an embodiment of the present application provides a voice interaction device, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method described in the first aspect or the third aspect.

[0033] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method described in the first aspect or the third aspect.

[0034] The voice interaction method provided by the present application combines pronunciation evaluation technology with voice synthesis technology, and can score the first audio data on the server, and obtain feedback voice of a specified timbre corresponding to the score. Among them, the speech synthesis model is trained using user-specified training data, so that the present application can generate feedback voice of any timbre according to the needs of the user, realizes the personalized customization of the feedback voice in the voice interaction method, and improves the selectivity of the feedback voice.

[0035] For example, a learning machine used by a certain user adopts the voice interaction method provided by this application. The user is more sensitive to his mother's voice, and to a certain extent, his mother's voice can improve the user's learning enthusiasm and concentration. Before the user uses the learning machine for pronunciation evaluation, his mother uploads her second audio data for the server to train the speech synthesis model. The trained speech synthesis model will be able to synthesize a voice that is very close to the mother's timbre. When the user uses the learning machine for pronunciation evaluation, the learning machine can upload the collected first audio data to the server for scoring, and the server determines the speech feedback text corresponding to the user's score; after that, the server calls the speech synthesis model to synthesize the feedback voice with the mother's timbre corresponding to the speech feedback text. Furthermore, the user can hear the feedback voice with the mother's timbre after the pronunciation evaluation. In other words, this application realizes the dialogue interaction between the user and a specific person during the pronunciation evaluation process, which improves the fun of the pronunciation evaluation product. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0037] Figure 1 A flow chart of a voice interaction method provided in an embodiment of the present application;

[0038] Figure 2 An interaction sequence diagram between a user terminal and a server in a pronunciation evaluation phase provided in an embodiment of the present application;

[0039] Figure 3 An interaction timing diagram between a user terminal and a server in a speech synthesis model training phase provided in an embodiment of the present application;

[0040] Figure 4 A schematic diagram of an audio acquisition process provided in an embodiment of the present application;

[0041] Figure 5 A schematic diagram of an audio preprocessing process provided in an embodiment of the present application;

[0042] Figure 6 A schematic diagram of a speech synthesis model training and generation process provided in an embodiment of the present application;

[0043] Figure 7 A schematic diagram of a network structure of a speech synthesis model provided in an embodiment of the present application;

[0044] Figure 8A schematic diagram of the structure of a voice interaction device provided in an embodiment of the present application;

[0045] Fig. 9 A schematic diagram of the structure of a voice interaction server provided in an embodiment of the present application;

[0046] Fig.10 A schematic diagram of the structure of a voice interaction user terminal provided in an embodiment of the present application;

[0047] Fig.11 A schematic diagram of the structure of another voice interaction device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0048] With the continuous development of science and technology, voice-based intelligent interaction has gradually spread to various fields. Pronunciation evaluation is a technology that automatically scores users' oral expressions, which can be used to determine whether the user's oral expression is standard. Specifically, pronunciation evaluation technology mainly uses computer intelligent algorithms to evaluate learners' pronunciation and diagnose pronunciation problems. For example, in the field of education, products with pronunciation evaluation functions can be used to evaluate users' pronunciation in various languages ​​such as Chinese and English to help users improve their oral expression skills. Nowadays, pronunciation evaluation technology is recognized by more and more people. It can evaluate oral fluency, completeness, accuracy and other dimensions, and can also be combined with question design to make learning more interesting.

[0049] Usually, various pronunciation evaluation products can first collect the user's pronunciation audio data, and then give a comprehensive score based on pronunciation evaluation technology from the dimensions of pronunciation accuracy, expression fluency, etc. When the user's score is low, the specific pronunciation error location can be pointed out to help the user improve oral expression. For example, after obtaining the pronunciation evaluation score, a certain brand of learning machine can feedback the specific score to the user through voice, and can also give corresponding praise or encouragement. For example, when a high score is obtained, the voice prompt is "Great!"; when a lower score is obtained, the voice prompt is "Try harder, you can do it." In this way, a natural human-computer interaction is formed, which can enhance the user's stickiness to the product.

[0050] Among them, the voice used for feedback can be a voice recorded in advance according to the specified content, or a voice synthesized in real time through speech synthesis technology. For the feedback voice recorded in advance, its timbre is the timbre of the recorder, which is fixed and single. Compared with the feedback voice synthesized based on speech synthesis technology, it is more flexible in text content, that is, it can support diverse text content, and then synthesize feedback voice corresponding to the text content based on a certain timbre. It can be seen that the feedback voice of pronunciation evaluation products usually uses only one or more fixed timbres, and the selectivity is low.

[0051] Based on this, the present application provides a voice interaction method, device and storage medium, which realizes the personalized customization of feedback voice in the voice interaction method by combining pronunciation evaluation technology and speech synthesis technology, and improves the selectivity of feedback voice.

[0052] In order to make the objectives, technical solutions and advantages of the present application more clear, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0053] It should be clear that the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.

[0054] In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances. In addition, in the description of the present application, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.

[0055] In one embodiment, see Figure 1 The voice interaction method provided in this application may include the following steps:

[0056] S102: Acquire first audio data.

[0057] See also Figure 2 , when the user-end device performs pronunciation evaluation on the user, the user-end device can obtain the user's audio data. For ease of description, the audio data used for pronunciation evaluation can be referred to as the first audio data. Afterwards, the user-end device can upload the first audio data to the server. Specifically, when the user-end device performs pronunciation evaluation on the user, it can collect the first audio data through a microphone or other sound pickup device, and upload the obtained first audio data to the online server, so as to utilize the server's strong computing power to perform pronunciation evaluation on the first audio data and generate a pronunciation evaluation result.

[0058] Among them, the user terminal device can be any terminal device such as a mobile phone, a laptop computer, etc. The sound pickup device can be integrated on the user terminal device, or it can be other devices connected to the user terminal device by wire or wirelessly. The server can be a single network device, or it can be one or more actual network devices in a server cluster. The above-mentioned server may include a processor, a memory, and a transceiver. The processor can be used to perform relevant processing of the voice interaction method, the memory can be used to store the data required and generated in the following processing process, and the transceiver can be used to receive and send relevant data in the following processing process.

[0059] S104: Obtain a pronunciation score of the first audio data and a voice feedback text corresponding to the pronunciation score.

[0060] In implementation, the server can obtain the pronunciation score of the first audio data and the voice feedback text corresponding to the pronunciation score. Specifically, after receiving the first audio data uploaded by the user end, the online server can call the pronunciation evaluation model to score the first audio data, generate the pronunciation score of the first audio data and the voice feedback text corresponding to the score.

[0061] In one embodiment, the pronunciation evaluation model may only output the pronunciation score of the first audio data. A plurality of voice feedback texts and a mapping relationship between each voice feedback text and a score may be pre-stored on the server. After obtaining the pronunciation score of the first audio data, the server may search for the corresponding voice feedback text in the mapping relationship table.

[0062] It should be noted that the pronunciation evaluation model can be built using network architectures such as recurrent neural networks (RNN) and convolutional neural networks (CNN). During actual construction, it is necessary to select appropriate model structures and parameters based on specific needs and data characteristics, and this application does not impose any restrictions on this.

[0063] In one embodiment, before the online server calls the pronunciation evaluation model to score the first audio data, the original first audio data may be preprocessed. For example, operations such as noise removal, standardization, and framing may be performed. Then, feature extraction is performed on the preprocessed first audio data. These features may include phonemes, acoustic features, etc. Among them, for different model structures, the method of feature extraction may also be different. Afterwards, the pronunciation evaluation model is called to use the features extracted in the previous stage to perform pronunciation scoring on the new audio signal.

[0064] S106, calling a speech synthesis model to synthesize a speech with a specified timbre corresponding to the speech feedback text; wherein the speech synthesis model is trained using the second audio data of the specified speaker.

[0065] In implementation, for the convenience of description, the audio data used to train the speech synthesis model may be referred to as the second audio data. Figure 2 The server can call a speech synthesis model to synthesize a speech with a specified timbre corresponding to the speech feedback text; wherein the speech synthesis model is trained using the second audio data of the specified speaker.

[0066] It is worth mentioning that before the user conducts the pronunciation evaluation based on the user terminal device, the user can specify the voice of interest, that is, select the designated speaker for feedback of the pronunciation evaluation result. Figure 3 , the user terminal device can obtain the second audio data of the specified speaker and send it to the server, so that the server uses the second audio data of the specified speaker as training data to train the speech synthesis model. Furthermore, the server can generate feedback speech with a specified timbre based on the trained speech synthesis model. Among them, the device used to collect audio data for user pronunciation evaluation and the device used to collect training data can be the same user terminal device or different user terminal devices, and this application does not limit this.

[0067] In one embodiment, the user can specify the timbre of interest to him / her on the human-computer interaction interface of the user terminal device, that is, select a designated speaker. If the user selects another person as the designated speaker, the second audio data of the designated speaker can be collected through the user terminal device, and the second audio data can be used as training data for the speech synthesis model; if the user selects the pronunciation evaluator himself / herself as the designated speaker, there is no need to collect the second audio data of the designated speaker, and the first audio data during the pronunciation evaluation can be used as training data for the speech synthesis model.

[0068] Correspondingly, when the user selects another person as the designated speaker, before the server uses the second audio data to train the speech synthesis model, the voice interaction method provided in the present application may also include: determining whether the designated speaker is a pronunciation evaluator; if the designated speaker is not the pronunciation evaluator, collecting the second audio data recorded by the designated speaker; the second audio data corresponds to the recording text sent by the server; and sending the second audio data to the server to train the speech synthesis model.

[0069] In implementation, the server may prepare the recording text in advance and send it to the user terminal device, so that the designated speaker can record the second audio data by referring to the recording text before the user performs the pronunciation assessment.

[0070] In one embodiment, the number of recorded texts prepared in advance can be 50 to 100 sentences. Taking Chinese text as an example, the design of the text needs to cover all Chinese pronunciation phonemes, including initials and finals with tones, a total of about 210. Taking the sentence "Children like to observe the world." as an example, its pinyin sequence is "hai2 zi5xi3 huan1 guan1 cha2 shi4jie4", where the pinyin can be separated by spaces, and the numbers can represent the tones, 1 to 4 correspond to the first to fourth tones, and 5 represents the light tone. The pinyin is split into the form of initials and finals, and the phoneme sequence "h ai2 zi5 x i3 h uan1 g uan1 ch a2sh i4 j ie4" can be obtained, which contains "g, h, j, x, ch, sh, z, a2, i3, i4, i5, ai2, ie4, uan1", a total of 14 phonemes. By constructing different sentences, the pronunciation phonemes are covered as comprehensively as possible.

[0071] After the recording text is prepared, the user terminal device can display the recording text to the designated speaker, collect the second audio data of the designated speaker through the sound pickup device, and then send the second audio data of the designated speaker to the server. Correspondingly, before step 106, the server can receive the second audio data of the designated speaker sent by the user terminal, and train the speech synthesis model based on the second audio data.

[0072] In one embodiment, in order to determine the training data type of the speech synthesis model, the voice interaction method provided in the present application may also include: determining whether the designated speaker is the pronunciation evaluator; if the designated speaker is the pronunciation evaluator, using the first audio data as the second audio data to train the speech synthesis model.

[0073] In practice, when users need to customize their own timbre, they can directly use the first audio data of the pronunciation assessment to train the speech synthesis model. This saves time in preparing recorded text and recording audio, and simplifies the personalized customization process of voice feedback. In addition, the longer the user uses the pronunciation assessment function, the more training data is used for speech synthesis, which can improve the performance of the speech synthesis model.

[0074] Furthermore, if the designated speaker is not the pronunciation evaluator, the server may train the speech synthesis model based on the second audio data recorded in advance by the designated speaker; the second audio data corresponds to the recording text sent by the server.

[0075] In implementation, the user terminal device may send a user instruction for indicating a designated speaker to the server, and the server may determine whether the designated speaker is the pronunciation evaluator according to the user instruction. The server may also determine whether the designated speaker is the pronunciation evaluator according to whether the second audio data of the designated speaker is received. Of course, the server may also use other methods to determine whether the designated speaker is the pronunciation evaluator, and this application does not limit this.

[0076] Among them, the way in which the server identifies whether the received data is the second audio data of the designated speaker can be determined according to the actual application, and this application does not limit this. For example, when the user terminal device sends the second audio data of the designated speaker, it can add special metadata or a specific flag bit in the data packet, and the server can determine whether the second audio data of the designated speaker is received based on whether the data packet carries the above data mark. For another example, the second audio data used for training the model can be encapsulated in a specified format, and the server can determine whether the second audio data of the designated speaker is received based on the data format.

[0077] In one embodiment, taking the user terminal device as a mobile phone as an example, the designated speaker can record the second audio data by following the instructions and referring to the recording text through the audio collection app or applet preset in the mobile phone. In order to improve the recording effect, the user terminal device can send a prompt message to the user suggesting wearing headphones for recording.

[0078] In one embodiment, the difference in recording environment will affect the recording effect of audio data. Therefore, before obtaining the second audio data, the voice interaction method provided by the present application may also include: detecting the audio recording environment to determine whether the audio recording environment meets the preset environmental conditions; if the audio recording environment meets the preset environmental conditions, then starting to collect the second audio data.

[0079] In implementation, see Figure 4 Before starting to record the second audio data, the user-end device can first record a specified length of ambient sound, such as 3s. Calculate the signal-to-noise ratio of the ambient sound. To a certain extent, the higher the signal-to-noise ratio, the better the sound quality. Therefore, the preset environmental condition can be a preset signal-to-noise ratio threshold, which is used to determine whether the recording environment is suitable. For example, if the signal-to-noise ratio threshold is set to 20dB, and the signal-to-noise ratio of the ambient sound is not higher than 20dB, the recording environment is more suitable. If the signal-to-noise ratio threshold is exceeded, a quieter recording space needs to be changed. Based on the recording environment detection processing, a better audio data recording effect can be guaranteed.

[0080] In one embodiment, the difference in recording equipment can also affect the recording effect of audio data. Therefore, before starting to collect the second audio data, the voice interaction method provided by the present application can also include: playing a test audio, collecting an audio segment of the test audio, and determining whether the audio segment meets the preset sound pickup condition; if the audio segment meets the preset sound pickup condition, then starting to collect the second audio data; otherwise, replacing the recording equipment of the user end.

[0081] In implementation, after meeting the recording environment requirements, the user-end device can also provide a test audio to determine whether the sound pickup is normal. During the recording of the test audio, the signal-to-noise ratio of the audio clip can be calculated in real time. If the signal-to-noise ratio of the test audio is continuously lower than 20dB, it means that the sound may not be picked up normally, and the user can be prompted to change the device until the recording environment test is passed. In this way, by performing environmental testing on the sound pickup device, a better audio data recording effect can be further guaranteed.

[0082] In another embodiment, the sampling rate of the test audio may also be counted. Generally, the higher the sampling rate, the more accurate the restoration of the original audio. The sampling rate is generally required to be no less than 16KHz.

[0083] After the above conditions are met, the user-end device can start to give out the recorded text one by one. After each recorded text is recorded, it can be determined whether the recorded content is consistent with the provided recorded text. For example, the audio data can be recognized by the speech recognition model, and the obtained recognition result can be compared with the pronunciation of the preset recorded text word by word. As long as the pronunciation is consistent, it is correct. Then, the proportion of correct words can be calculated. If the correct proportion exceeds the preset threshold, the consistency test is passed, otherwise the recorded text is re-recorded. Save the second audio data that passes the consistency test, and then continue to record the subsequent recorded text until all the recorded texts are recorded.

[0084] In implementation, if the user needs to customize his or her own voice for feedback on the pronunciation evaluation results, that is, when the pronunciation evaluator is the same as the designated speaker, the additional audio collection process can be omitted. Because during the pronunciation evaluation, in order to evaluate the quality of the pronunciation, it is necessary to record the user's pronunciation audio, and these audio data used for pronunciation evaluation can be directly used as training data for the speech synthesis model, eliminating the need to collect audio data for training the speech synthesis model. In this way, the time for preparing the recorded text is saved and the personalized customization process of voice feedback is simplified. Moreover, the longer the user uses the pronunciation evaluation function, the more data is used to train the speech synthesis model, which can further improve the performance of the speech synthesis model.

[0085] In one embodiment, due to factors such as the model and configuration of the device, as well as the recording environment, the quality of the collected original audio data may vary. In order to improve the quality of the audio signal, before using the second audio data to train the speech synthesis model, or before performing pronunciation evaluation on the first audio data, the present application may also perform one or more of the following preprocessing on the second audio data and / or the first audio data: uniformly set the sampling rate of each audio data to a specified sampling rate; perform noise reduction processing on the audio data; perform amplitude normalization processing on the audio data; set the first and last silent segments of the audio data to a specified length.

[0086] In practice, the higher the sampling rate of audio data, the higher the quality of the sampling, and the better it can restore the details of the original audio signal. At the same time, as the sampling rate increases, the size of the audio file will also increase accordingly, requiring more storage space and computing resources. Therefore, in order to meet the training requirements of the speech synthesis model and reduce the cost of model training, a suitable sampling rate can be preset for the model's training data, such as 16kHz.

[0087] For details, see Figure 5 Since the sampling rates of audio data collected by different devices may be different, the online server can adjust the audio sampling rate to a preset sampling rate after obtaining the audio data. For example, the audio sampling rate of a certain brand of mobile phone is 48kHz. After the second audio data is collected using the mobile phone, the sampling rate of the second audio data can be reduced to 16kHz.

[0088] In order to remove interference and noise in the audio data, the audio is processed through the noise reduction model to filter out the background noise.

[0089] Furthermore, in order to keep the volume of the audio data in a stable range as much as possible, the amplitude of the denoised audio can be normalized to ensure that the amplitude of all audio is relatively uniform.

[0090] Finally, in order to ensure the stability of the audio signal, the original silent segments at the beginning and end of the audio can be deleted, and silent segments of a specified length can be spliced. For example, a 300ms silent segment can be spliced ​​at the beginning and end of the audio data after deleting the silent segment, respectively, to ensure that the lengths of the silent segments at the beginning and end of all audios are close. Among them, based on the characteristic that the sound energy of the silent segment is lower than that of the speech segment, the signal-to-noise ratio can be used to determine how long the original silent segments at the beginning and end of the audio data are.

[0091] Based on the above preprocessing steps, the quality of the audio signal can be improved, and more accurate training data can be obtained, so that the speech synthesis model can better learn the audio data during training.

[0092] In one embodiment, the sampling rate of the original audio data is the same as the sampling rate preset during model training, and therefore, the preprocessing step of adjusting the audio sampling rate may not be performed on the original audio data.

[0093] In one embodiment, the second audio data may be firstly subjected to feature extraction, and then the speech synthesis model may be trained using the extracted features. Accordingly, the step of training the speech synthesis model using the second audio data of a designated speaker may specifically include: extracting feature data of the second audio data and the recording text corresponding to the second audio data; and fine-tuning the pre-trained speech synthesis model based on the feature data.

[0094] In practice, audio data often contains many dimensions. For model training, not all data dimensions are useful. Feature extraction of audio data used to train speech synthesis models can reduce the dimensionality and complexity of the data, making model training more efficient and accurate.

[0095] In one embodiment, the type of features to be extracted can be determined based on actual needs. For example, extracting feature data of the second audio data and the recording text corresponding to the second audio data may specifically include: aligning the second audio data and the recording text using an alignment model to obtain the duration features of each phoneme in the second audio data; extracting frequency domain features from the second audio data; extracting phoneme-level acoustic features from the second audio data based on the duration features; wherein the acoustic features include fundamental frequency features and / or energy features.

[0096] In implementation, see Figure 6 Before training the speech synthesis model, the features of audio and text can be extracted. First, use an alignment model, such as the MFA model (Montreal-Forced-Aligner), to align the preprocessed audio data with the corresponding recorded text. Take "Children like to observe the world." as an example. Its phoneme sequence is "sil h ai2 zi5 x i3 h uan1 g uan1 ch a2 sh i4 j ie4 sil", which is recorded as the text feature. Among them, "sil" represents the silence segment at the beginning and end of the audio. Assuming that the duration of the audio data is 2s, after the alignment model, the estimated duration corresponding to each phoneme in the phoneme sequence will be obtained. For example, the lengths of the first three phonemes "sil", "h" and "ai2" are 0.3s, 0.1s and 0.2s respectively. This feature is recorded as the duration feature.

[0097] Furthermore, frequency domain features may be extracted from the audio data. For example, the audio data may be subjected to Fourier transform to obtain frequency domain features. The frequency domain features may include one or more of the features such as amplitude spectrum, phase spectrum, Fourier coefficients, etc.

[0098] In addition, the fundamental frequency and energy features corresponding to each phoneme can be obtained by combining the duration features. For example, the audio data can be first subjected to Fourier transformation to obtain the amplitude, and then the square of the amplitude can be calculated to obtain the energy feature. For another example, an open source library, such as pyworld, can be called to obtain the fundamental frequency features of the audio data.

[0099] After extracting the above features, you can start fine-tuning the personalized speech synthesis model. Since the amount of recorded audio data is usually set to 50 to 100, you can train a basic speech synthesis model (or pre-trained speech synthesis model) in advance, and then use this batch of audio data to fine-tune the basic speech synthesis model. The fine-tuning training process usually takes 30 minutes to 1 hour to generate a speech synthesis model with good performance that can generate the same timbre as the recorder, and the development cycle is short.

[0100] It is worth mentioning that in actual use, both feature extraction and model training can be completed through automated scripts without manual intervention, which greatly improves the efficiency of model production and reduces labor costs. Among them, these automated scripts can be written in programming languages ​​such as Python, R or Scala, and this application does not limit this.

[0101] In one embodiment, see Figure 7 The network structure of the speech synthesis model mainly includes: text feature encoder + acoustic feature adapter + duration predictor + decoder + vocoder, which can be built by deep neural networks such as Transfomer, CNN, Generative Adversarial Networks (GAN), and this application does not limit this. During training, the input data of the text feature encoder is a phoneme sequence, the acoustic feature adapter can include a fundamental frequency predictor and an energy predictor, and its output requires fundamental frequency and energy features as references to calculate loss (model loss), the output of the duration predictor requires duration features as references to calculate loss, the output of the decoder requires frequency domain features as references to calculate loss, and the output of the vocoder is synthesized audio, that is, the audio data of the specified speaker.

[0102] It is worth mentioning that the structure of the fine-tuned speech synthesis model is the same as that of the pre-trained speech synthesis model, but the training process is slightly different. For example, during the pre-training process, the speech synthesis model needs to start learning from randomly initialized parameters, and the training time is relatively long. During the fine-tuning training process, the speech synthesis model only needs to make minor adjustments based on the parameters of the pre-trained model to adapt to specific tasks, and the training time is relatively short. It can be seen that this application is based on the pre-trained speech synthesis model, and fine-tuning the model using the audio data of the specified speaker, which can significantly save training time. Among them, the data used for pre-training the model can usually be obtained through channels such as data companies and open source databases, and this application does not impose any restrictions on this.

[0103] S108, feeding back the designated timbre voice.

[0104] In practice, based on the above training process, the server will update the model parameters to obtain a speech synthesis model that can generate speech with the same or similar timbre as the specified speaker. The user can first select the specified timbre through the user terminal device, and then after the pronunciation evaluation and scoring is completed, the speech synthesis online service can be called normally, that is, the speech synthesis model is called to synthesize the feedback speech with the specified timbre.

[0105] S110, playing a voice with a specified timbre.

[0106] In implementation, after receiving the specified timbre voice returned by the server, the user-end device can play the specified timbre voice, so that the pronunciation evaluator can hear the voice feedback of the specified timbre for his pronunciation evaluation results, thereby realizing personalized voice interaction.

[0107] In one embodiment, after step S110, the voice interaction method provided by the present application may further include: receiving interactive feedback sent by the user terminal; if the interactive feedback indicates that the pronunciation effect of the specified timbre voice is not good, supplementing the recording text for the designated speaker to re-record the second audio data for re-training the speech synthesis model.

[0108] In implementation, when the designated speaker is the same as the pronunciation evaluator, the first audio data collected during the pronunciation evaluation process is the training data for the speech synthesis model. Since the number of sentences used in the pronunciation evaluation is limited, if it cannot cover all pronunciation factors, the training effect will not be ideal after the speech synthesis model is trained using the first audio data recorded during the pronunciation evaluation process. Based on this, the present application can also provide a user feedback channel to collect user opinions on the specified timbre voice. For example, after the user hears the specified timbre voice played by the user-end device, the user can mark which words and sentences in the specified timbre voice have poor pronunciation effects on the user-end device. Afterwards, the user-end device can send the user's interactive feedback to the server, and the server supplements the recorded text corresponding to the words and sentences marked by the user, and specifically allows the user to re-record some audio to train the speech synthesis model again. In this way, more training phonemes can be further covered, and the training effect of the speech synthesis model can be improved.

[0109] This application combines pronunciation evaluation technology with speech synthesis technology, and can score the first audio data on the server, and obtain feedback speech of a specified timbre corresponding to the score. Among them, the speech synthesis model is trained using user-specified training data, so that this application can generate feedback speech of any timbre according to the user's needs, realize the personalized customization of feedback speech in the speech interaction method, and improve the selectivity of feedback speech. In addition, this application realizes the dialogue interaction between the user and the specific person during the pronunciation evaluation process, which enhances the fun of the pronunciation evaluation product.

[0110] For example, when the pronunciation evaluation user is a child, the child's parents or any other parent can record several audios in advance to generate a speech synthesis model similar to their voice. When providing voice feedback on the evaluation results of the child's pronunciation, the parent's voice can be selected to make the child feel that they are accompanied by family members and feel more intimate. Similarly, when the product target group is other types of users, users can also be supported to specify the voice they are interested in.

[0111] It is worth mentioning that in each of the above embodiments, the main load of the user terminal device is to collect audio data and play feedback voice, and the more complex pronunciation evaluation steps and voice synthesis steps can be executed by the online server. In this way, the computing load of the user terminal device can be significantly reduced. Of course, if the computing conditions permit, Figure 1 Each of the steps shown can also be performed by a single voice interaction device. For example, the user terminal device can collect the first audio data, call the pronunciation evaluation model to score the first audio data, and generate a voice feedback text corresponding to the pronunciation score. Then, the user terminal device calls the trained speech synthesis model to synthesize the corresponding specified timbre voice, and uses a speaker to play the specified timbre voice to achieve voice interaction with the user.

[0112] Correspondingly, based on the same technical concept, the embodiment of the present application also provides a voice interaction device. The specific implementation process can be found in the above method embodiment, which will not be repeated here. Figure 8 , the voice interaction device may include:

[0113] An audio acquisition module, used to acquire first audio data;

[0114] A pronunciation evaluation module, configured to call a pronunciation evaluation model to score the first audio data, obtain a pronunciation score of the user, and determine a voice feedback text corresponding to the pronunciation score;

[0115] A speech synthesis module, used for calling a speech synthesis model to synthesize a speech of a specified timbre corresponding to the speech feedback text; wherein the speech synthesis model is trained using the second audio data of the specified speaker;

[0116] The voice playing module is used to play the voice with the specified timbre.

[0117] Based on the same technical concept, the embodiment of the present application also provides a voice interaction server. The specific implementation process can be found in the above method embodiment, which will not be repeated here. Fig. 9 , the voice interaction server may include:

[0118] A pronunciation evaluation module, used to obtain a pronunciation score of the first audio data and a voice feedback text corresponding to the pronunciation score;

[0119] A speech synthesis module, used for calling a speech synthesis model to synthesize a speech of a specified timbre corresponding to the speech feedback text; wherein the speech synthesis model is trained using the second audio data of the specified speaker;

[0120] The voice feedback module is used to feed back the specified timbre voice to the user terminal so that the user terminal plays the specified timbre voice.

[0121] Furthermore, the voice interaction server may also include:

[0122] The model training module is used to receive the second audio data of the designated speaker sent by the user terminal, and train the speech synthesis model based on the second audio data.

[0123] Furthermore, the model training module can be specifically used for:

[0124] Determining whether the designated speaker is a pronunciation evaluator;

[0125] If the designated speaker is the pronunciation evaluator, the first audio data is used as the second audio data to train the speech synthesis model.

[0126] Furthermore, the model training module can also be specifically used for:

[0127] If the designated speaker is not the pronunciation evaluator, the speech synthesis model is trained based on the second audio data recorded in advance by the designated speaker; the second audio data corresponds to the recording text sent by the server.

[0128] Furthermore, the voice interaction server may also include an interactive feedback module for:

[0129] Receiving interactive feedback sent by the user terminal;

[0130] If the interactive feedback indicates that the pronunciation of the specified timbre voice is not good, the recording text is supplemented so that the specified speaker can record the second audio data for re-training the speech synthesis model.

[0131] Furthermore, the voice interaction server may further include a preprocessing module, configured to perform one or more of the following preprocessing on the second audio data:

[0132] The sampling rates of the second audio data are uniformly set to a specified sampling rate;

[0133] Performing noise reduction processing on the second audio data;

[0134] Performing amplitude normalization processing on the second audio data;

[0135] The first and last silent segments of the second audio data are set to a specified length.

[0136] Furthermore, the model training module can also be used for:

[0137] Extracting feature data of the second audio data and a recording text corresponding to the second audio data;

[0138] The pre-trained speech synthesis model is fine-tuned based on the feature data.

[0139] Furthermore, the model training module can also be specifically used for:

[0140] Aligning the second audio data with the recorded text using an alignment model to obtain a duration feature of each phoneme in the second audio data;

[0141] extracting frequency domain features from the second audio data;

[0142] Extracting phoneme-level acoustic features from the second audio data based on the duration features; wherein the acoustic features include fundamental frequency features and / or energy features.

[0143] Based on the same technical concept, the embodiment of the present application also provides a voice interactive user terminal. The specific implementation process can be found in the above method embodiment, which will not be repeated here. Fig.10 , the voice interaction client may include:

[0144] An audio acquisition module, used to acquire first audio data;

[0145] An audio uploading module, used for uploading the first audio data to a server;

[0146] A voice playing module, used for receiving and playing the voice with a specified timbre fed back by the server;

[0147] The specified timbre voice is generated by the server using the following steps:

[0148] Calling a pronunciation evaluation model to score the first audio data to obtain a pronunciation score of the user, and determining a voice feedback text corresponding to the pronunciation score;

[0149] Calling a speech synthesis model to synthesize a speech with a specified timbre corresponding to the speech feedback text; wherein,

[0150] The speech synthesis model is trained using second audio data of a designated speaker.

[0151] Furthermore, the audio acquisition module can be specifically used for:

[0152] Determining whether the designated speaker is a pronunciation evaluator;

[0153] If the designated speaker is not the pronunciation evaluator, collecting second audio data recorded by the designated speaker; the second audio data corresponds to the recording text sent by the server;

[0154] The second audio data is sent to the server to train the speech synthesis model.

[0155] Furthermore, the audio acquisition module can also be used for:

[0156] Detecting the audio recording environment to determine whether the audio recording environment meets preset environmental conditions;

[0157] If the audio recording environment meets the preset environmental condition, the second audio data is started to be collected.

[0158] Furthermore, the audio acquisition module can also be used for:

[0159] Playing a test audio, collecting an audio segment of the test audio, and determining whether the audio segment meets a preset sound pickup condition;

[0160] If the audio segment meets the preset sound pickup condition, then start collecting the second audio data; otherwise, replace the recording device at the user end.

[0161] Based on the same technical concept, the present application embodiment also provides a voice interaction device. The specific implementation process can be found in the above method embodiment, which will not be repeated here. Fig.11 The voice interaction device may include: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the processing performed by the user end or the server in any of the above-mentioned voice interaction methods.

[0162] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware. Based on this understanding, the above technical solution can essentially or contribute to the prior art in the form of a software product, and the voice interactive software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including storing a number of instructions to enable an electronic device to execute the methods described in each embodiment or some parts of the embodiments.

[0163] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A voice interaction method, characterized in that: The method comprises: Obtaining a pronunciation score of the first audio data and a voice feedback text corresponding to the pronunciation score; Calling a speech synthesis model to synthesize a speech with a specified timbre corresponding to the speech feedback text; wherein the speech synthesis model is trained using second audio data of a specified speaker; Feedback the specified timbre voice to the user terminal so that the user terminal plays the specified timbre voice.

2. The voice interaction method according to claim 1, characterized in that: Before calling the speech synthesis model to synthesize the specified timbre speech corresponding to the speech feedback text, the method further includes: Receive second audio data of the designated speaker sent by the user terminal, and train the speech synthesis model based on the second audio data.

3. The voice interaction method according to claim 2, characterized in that: The training of the speech synthesis model based on the second audio data specifically includes: Determining whether the designated speaker is a pronunciation evaluator; If the designated speaker is the pronunciation evaluator, the first audio data is used as the second audio data to train the speech synthesis model.

4. The voice interaction method according to claim 3, characterized in that: If the designated speaker is not the pronunciation evaluator, the speech synthesis model is trained based on the second audio data recorded in advance by the designated speaker; the second audio data corresponds to the recording text sent by the server.

5. The voice interaction method according to claim 1, wherein: After the user terminal plays the specified timbre voice, the method further includes: Receiving interactive feedback sent by the user terminal; If the interactive feedback indicates that the pronunciation of the specified timbre voice is not good, the recording text is supplemented so that the specified speaker can record the second audio data for re-training the speech synthesis model.

6. The voice interaction method according to claim 1, wherein: Before training the speech synthesis model, the method further includes: Perform one or more of the following preprocessing on the second audio data: The sampling rates of the second audio data are uniformly set to a specified sampling rate; Performing noise reduction processing on the second audio data; Performing amplitude normalization processing on the second audio data; The first and last silent segments of the second audio data are set to a specified length.

7. The voice interaction method according to claim 1, wherein: The step of using the second audio data of the designated speaker to train the speech synthesis model specifically includes: Extracting feature data of the second audio data and a recording text corresponding to the second audio data; The pre-trained speech synthesis model is fine-tuned based on the feature data.

8. The voice interaction method according to claim 7, characterized in that: The step of extracting the feature data of the second audio data and the audio recording text corresponding to the second audio data specifically includes: Aligning the second audio data with the recorded text using an alignment model to obtain a duration feature of each phoneme in the second audio data; extracting frequency domain features from the second audio data; Extracting phoneme-level acoustic features from the second audio data based on the duration features; wherein the acoustic features include fundamental frequency features and / or energy features.

9. A voice interaction server, characterized in that: include: A pronunciation evaluation module, used to obtain a pronunciation score of the first audio data and a voice feedback text corresponding to the pronunciation score; A speech synthesis module, used for calling a speech synthesis model to synthesize a speech of a specified timbre corresponding to the speech feedback text; wherein the speech synthesis model is trained using the second audio data of the specified speaker; The voice feedback module is used to feed back the specified timbre voice to the user terminal so that the user terminal plays the specified timbre voice.

10. A voice interaction method, characterized in that: The method comprises: Acquire first audio data, and upload the first audio data to a server; Receive and play the specified tone voice fed back by the server; The specified timbre voice is generated by the server using the following steps: Calling a pronunciation evaluation model to score the first audio data to obtain a pronunciation score of the user, and determining a voice feedback text corresponding to the pronunciation score; A speech synthesis model is called to synthesize a speech with a specified timbre corresponding to the speech feedback text; wherein the speech synthesis model is trained using the second audio data of a specified speaker.

11. The voice interaction method according to claim 10, wherein: Before the server uses the second audio data to train the speech synthesis model, the method further includes: Determining whether the designated speaker is a pronunciation evaluator; If the designated speaker is not the pronunciation evaluator, collecting second audio data recorded by the designated speaker; the second audio data corresponds to the recording text sent by the server; The second audio data is sent to the server to train the speech synthesis model.

12. The voice interaction method according to claim 10, wherein: Before the server uses the second audio data to train the speech synthesis model, the method further includes: Detecting the audio recording environment to determine whether the audio recording environment meets preset environmental conditions; If the audio recording environment meets the preset environmental condition, the second audio data is started to be collected.

13. The voice interaction method according to claim 10, wherein: Before the server uses the second audio data to train the speech synthesis model, the method further includes: Playing a test audio, collecting an audio segment of the test audio, and determining whether the audio segment meets a preset sound pickup condition; If the audio segment meets the preset sound pickup condition, then start collecting the second audio data; otherwise, replace the recording device at the user end.

14. A voice interaction user terminal, characterized in that: include: An audio acquisition module, used to acquire first audio data; An audio uploading module, used for uploading the first audio data to a server; A voice playing module, used for receiving and playing the voice with a specified timbre fed back by the server; The specified timbre voice is generated by the server using the following steps: Calling a pronunciation evaluation model to score the first audio data to obtain a pronunciation score of the user, and determining a voice feedback text corresponding to the pronunciation score; A speech synthesis model is called to synthesize a speech with a specified timbre corresponding to the speech feedback text; wherein the speech synthesis model is trained using the second audio data of a specified speaker.

15. A voice interaction device, characterized in that: The voice interaction device comprises: An audio acquisition module, used to acquire first audio data; A pronunciation evaluation module, configured to call a pronunciation evaluation model to score the first audio data, obtain a pronunciation score of the user, and determine a voice feedback text corresponding to the pronunciation score; A speech synthesis module, used for calling a speech synthesis model to synthesize a speech of a specified timbre corresponding to the speech feedback text; wherein the speech synthesis model is trained using the second audio data of the specified speaker; The voice playing module is used to play the voice with the specified timbre.

16. A voice interaction device, characterized in that: The voice interaction device includes: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method described in any one of claims 1-8 and 10-13.

17. A computer-readable storage medium, characterized in that: The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1-8 and 10-13.

Citation Information

Patent Citations

  • Computer system for assisting spoken language learning

    CN101551947A

  • Intelligent preschool education and learning system and method

    CN107798931A

  • Speech broadcasting method and device, electronic equipment and storage medium

    CN110600000A

  • Model training method, voice playing method and device and storage medium

    CN111816168A

  • Parent-child interaction type early education system

    CN112365752A