Acoustic Feature Prediction Model Training Method, Apparatus, Electronic Device, and Computer Readable Storage Medium
By imitating the way humans use the sounding method under Lombard effect, the acoustic feature prediction model is used to improve the clarity and nature of voice broadcasts in noisy home scenes, solving the problem that the voice broadcasts of smart devices cannot be accurately received by users in noisy environments.
Patent Information
- Application Number
- CN202110069952.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-19
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-01-19
AI Technical Summary
In noisy home scenes, voice broadcasts feedback from smart devices may not be accurately received by users, affecting the smoothness of human-computer interaction.
By imitating humans to actively change their voice-making methods under Lombard effect, using acoustic feature prediction model training method, predict the acoustic features corresponding to home scenes, and output the sound prediction features to be synthesized with Lombard speech acoustic style to improve the speech clarity and nature of the speech synthesis process.
It realizes improving the clarity and nature of voice broadcasts in noisy home scenes, allowing users to receive feedback from smart devices more accurately, and improving the smoothness of human-computer interaction.
Smart Images

Figure CN114822484B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method, device, electronic device, and computer-readable storage medium for training an acoustic feature prediction model. Background Art
[0002] In daily life, various noises fill the home environment, and people often need to communicate in noisy environments. French otolaryngologist Étienne Lombard discovered through research in 1909 that when communicating in a noisy environment, the speaker has to actively change the way of speaking to improve the sound effect in the hope that the other party can hear clearly. Through research, it has been found that even when the same person pronounces the same speech, the speech features are different in different environments. The changed features include increasing the pitch, tone, loudness, and formant features of the sound. This phenomenon is called the Lombard effect. With the rise of smart homes, how to improve the smoothness of human-computer interaction to better serve humans has become an issue of concern in this field.
[0003] Currently, electronic devices such as smart devices in the home environment can sense changes in the acoustic environment state and give reasonable answers to the user's questions. For example, when the smart device detects that the user issues an instruction message "turn on the microwave oven", it controls the microwave oven to turn on and voice broadcasts "the microwave oven has been turned on" to complete the response to the instruction message. However, through research, it has been found that the content of the voice broadcast by the smart device cannot be accurately received by the user in many cases. For example, the voice broadcast in a noisy environment is drowned out by the ambient sound, making it impossible for the user to clearly receive the broadcast voice, which affects the smoothness of human-computer interaction in the smart home scenario and cannot meet the actual application requirements. Summary of the Invention
[0004] Based on the above problems, the inventors analyzed and concluded that the prerequisite for the user to accurately receive reasonable content such as the voice broadcast of "the microwave oven has been turned on" is that the smart device can change the way of speaking and actively improve the clarity and naturalness of the synthesized voice in different acoustic environments. However, in existing research, insufficient consideration has been given to how the smart device synthesizes speech and actively changes the way of speaking like humans under the Lombard effect to improve speech clarity and allow the user to receive accurate information, resulting in the possibility that the voice broadcast feedback by the smart device may not be received by the user. In view of this, the inventors proposed to "imitate humans" in a noisy home environment and actively change the way of speaking under the Lombard effect to improve the clarity and naturalness of the voice broadcast (i.e., emit Lombard speech). One of the keys lies in predicting the acoustic features corresponding to the home environment.
[0005] One of the objectives of the present invention includes, for example, providing a method, apparatus, electronic device, and computer-readable storage medium for training an acoustic feature prediction model to achieve reliable prediction of acoustic features corresponding to a home scene.
[0006] Embodiments of the present invention may be implemented as follows:
[0007] In a first aspect, an embodiment of the present invention provides a method for training an acoustic feature prediction model, including:
[0008] Obtaining text information to be output and training style embedding information corresponding to a home scene;
[0009] Inputting the text information to be output and the training style embedding information into an acoustic feature prediction network to obtain predicted features of the sound to be synthesized corresponding to the home scene.
[0010] By inputting both the text information to be output and the training style embedding information into the acoustic feature prediction network, reliable prediction of acoustic features corresponding to the home scene can be achieved based on two dimensions of the text information to be output and the training style embedding information, and predicted features of the sound to be synthesized with a Lombard speech acoustic style can be output, thereby improving the matching degree between the synthesized speech and the home scene in the subsequent speech synthesis process.
[0011] In a second aspect, an embodiment of the present invention provides an apparatus for training an acoustic feature prediction model, including:
[0012] A data acquisition module for obtaining text information to be output and training style embedding information corresponding to a home scene;
[0013] An acoustic feature prediction model training module for inputting the text information to be output and the training style embedding information into an acoustic feature prediction network to obtain predicted features of the sound to be synthesized corresponding to the home scene.
[0014] By inputting both the text information to be output and the training style embedding information into the acoustic feature prediction network, reliable prediction of acoustic features corresponding to the home scene can be achieved based on two dimensions of the text information to be output and the training style embedding information, and predicted features of the sound to be synthesized with a Lombard speech acoustic style can be output, thereby improving the matching degree between the synthesized speech and the home scene in the subsequent speech synthesis process.
[0015] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the method for training an acoustic feature prediction model according to any one of the foregoing embodiments when executing the program. Correspondingly, the electronic device includes the beneficial effects in the method for training an acoustic feature prediction model.
[0016] Fourthly, an embodiment of the present invention provides a computer-readable storage medium, which includes a computer program. When the computer program runs, it controls the electronic device where the computer-readable storage medium is located to execute the acoustic feature prediction model training method according to any one of the foregoing embodiments. Correspondingly, this computer-readable storage medium includes the beneficial effects in the acoustic feature prediction model training method.
[0017] To make the above objects, features, and advantages of the embodiments of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0019] Figure 1 FIG. shows a schematic diagram of an application scenario provided by an embodiment of the present invention.
[0020] Figure 2 FIG. shows a schematic flowchart of an acoustic feature prediction model training method provided by an embodiment of the present invention.
[0021] Figure 3 FIG. shows a schematic diagram of the training architecture of an acoustic feature prediction model provided by an embodiment of the present invention.
[0022] Figure 4 FIG. shows a schematic flowchart of a training method for a scene acoustic style extractor provided by an embodiment of the present invention.
[0023] Figure 5 FIG. shows a schematic diagram of the training architecture of a scene acoustic style extractor provided by an embodiment of the present invention.
[0024] Figure 6 FIG. shows another schematic flowchart of a training method for a scene acoustic style extractor provided by an embodiment of the present invention.
[0025] Figure 7 FIG. shows another schematic diagram of the training architecture of a scene acoustic style extractor provided by an embodiment of the present invention.
[0026] Figure 8 FIG. shows another schematic flowchart of a training method for a scene acoustic style extractor provided by an embodiment of the present invention.
[0027] Figure 9 Shows another schematic diagram of the training architecture of a scene acoustic style extractor provided by an embodiment of the present invention.
[0028] Figure 10 Shows another schematic flowchart of a method for training a scene acoustic style extractor provided by an embodiment of the present invention.
[0029] Figure 11 Shows one of the schematic diagrams of a reference encoder provided by an embodiment of the present invention.
[0030] Figure 12 Shows another schematic diagram of a reference encoder provided by an embodiment of the present invention.
[0031] Figure 13 Shows yet another schematic diagram of a reference encoder provided by an embodiment of the present invention.
[0032] Figure 14 Shows another schematic diagram of the training architecture of a scene acoustic style extractor provided by an embodiment of the present invention.
[0033] Figure 15 Shows a schematic flowchart of a method for training a first scene classification model provided by an embodiment of the present invention.
[0034] Figure 16 Shows another schematic flowchart of a method for training a first scene classification model provided by an embodiment of the present invention.
[0035] Figure 17 Shows a schematic flowchart of a speech synthesis method provided by an embodiment of the present invention.
[0036] Figure 18 Shows a schematic diagram of the implementation principle of a speech synthesis method provided by an embodiment of the present invention.
[0037] Figure 19 Shows another schematic diagram of the implementation principle of a speech synthesis method provided by an embodiment of the present invention.
[0038] Figure 20 Shows a schematic diagram of the implementation principle of a scene acoustic style extractor provided by an embodiment of the present invention.
[0039] Figure 21 Shows another schematic diagram of the implementation principle of a scene acoustic style extractor provided by an embodiment of the present invention.
[0040] Figure 22 Shows a schematic diagram of the implementation principle of an acoustic feature prediction model provided by an embodiment of the present invention.
[0041] Figure 23Shows a schematic diagram of the implementation principle of a speech synthesis method provided by an embodiment of the present invention.
[0042] Figure 24 Shows a schematic diagram of the implementation principle of another method for synthesizing and outputting response speech data provided by an embodiment of the present invention.
[0043] Figure 25 Shows a schematic flowchart of a speech synthesis model training method provided by an embodiment of the present invention.
[0044] Figure 26 Shows a schematic diagram of the training architecture of a speech synthesis model provided by an embodiment of the present invention.
[0045] Figure 27 Shows an exemplary structural block diagram of a first speech synthesis device provided by an embodiment of the present invention.
[0046] Figure 28 Shows an exemplary structural block diagram of a second speech synthesis device provided by an embodiment of the present invention.
[0047] Figure 29 Shows an exemplary structural block diagram of a scene classification model training device provided by an embodiment of the present invention.
[0048] Figure 30 Shows an exemplary structural block diagram of a scene acoustic style extractor training device provided by an embodiment of the present invention.
[0049] Figure 31 Shows an exemplary structural block diagram of an acoustic feature prediction model training device provided by an embodiment of the present invention.
[0050] Figure 32 Shows an exemplary structural block diagram of a speech synthesis model training device provided by an embodiment of the present invention.
[0051] Icons: 100 - electronic device; 110 - memory; 120 - processor; 130 - communication module; 140 - first speech synthesis device; 141 - information determination module; 142 - response speech data synthesis module; 150 - second speech synthesis device; 151 - predicted acoustic feature acquisition module; 152 - information synthesis module; 160 - scene classification model training device; 161 - environmental acoustic feature acquisition module; 162 - scene classification network training module; 170 - scene acoustic style extractor training device; 171 - noisy environment acoustic feature acquisition module; 172 - scene acoustic style extractor training module; 180 - acoustic feature prediction model training device; 181 - data acquisition module; 182 - acoustic feature prediction model training module; 190 - speech synthesis model training device; 191 - information acquisition module; 192 - to-be-output speech data synthesis module. Detailed implementation manners
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0053] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0054] It should be noted that the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device.
[0055] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it will not be repeatedly defined and explained in the subsequent drawings without special instructions.
[0056] It should be noted that, without conflict, the features in the embodiments of the present invention can be combined with each other.
[0057] Please refer to Figure 1 , which is a block diagram of an electronic device 100 provided in this embodiment. The electronic device 100 in this embodiment can be various processing devices capable of information interaction and processing. For example, the electronic device 100 can be an intelligent device capable of voice interaction with users. For another example, the electronic device can be a server, a processing platform, etc. capable of model training and voice synthesis.
[0058] The electronic device 100 may include a memory 110, a processor 120, and a communication module 130. The elements of the memory 110, the processor 120, and the communication module 130 are directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines.
[0059] Among them, the memory 110 is used to store programs or data. The memory 110 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0060] The processor 120 is used to read / write the data or programs stored in the memory 110 and execute corresponding functions.
[0061] The communication module 130 is used to establish a communication connection between the electronic device 100 and other communication terminals through the network, and is used to transmit and receive data through the network.
[0062] It should be understood that Figure 1 The structure shown is only a schematic diagram of the structure of the electronic device 100. The electronic device 100 may further include more or fewer components than those shown Figure 1 in the figure, or have a different configuration from that shown Figure 1 in the figure. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof. For example, in the case where the electronic device 100 is an intelligent device capable of voice interaction with users, the electronic device 100 may further include a voice module, and the electronic device 100 can collect and output voices through the voice module.
[0063] Please refer to Figure 2 , which is a schematic flowchart of a method for training an acoustic feature prediction model provided by an embodiment of the present invention, and can be executed by the electronic device 100 shown Figure 1 in the figure, for example, can be executed by the processor 120 in the electronic device 100. The method for training an acoustic feature prediction model includes S110 and S120.
[0064] S110, obtaining the text information to be output and the training style embedding information corresponding to the home scene.
[0065] S120, inputting the text information to be output and the training style embedding information into an acoustic feature prediction network to obtain the predicted features of the sound to be synthesized corresponding to the home scene.
[0066] The training style embedding information corresponding to the home scene represents the scene acoustic style corresponding to the home scene.
[0067] In this embodiment, the training style embedding information corresponding to the home scene is skillfully introduced as a new dimension for voice synthesis consideration. Both the text information to be output and the training style embedding information are input into the acoustic feature prediction network to obtain the predicted features of the sound to be synthesized corresponding to the home scene. The predicted features of the sound to be synthesized have the Lombard speech acoustic style. Based on the two dimensions of the text information to be output and the training style embedding information, reliable prediction of the acoustic features corresponding to the home scene can be achieved, and the predicted acoustic features with the Lombard speech acoustic style are output.
[0068] In one implementation, the speech uttered by the user in a quiet scene can be collected as normal speech, and the acoustic feature prediction network is jointly trained using the normal speech and the training style embedding information to obtain the predicted features of the sound to be synthesized corresponding to the home scene. For example, when the training style embedding information is to increase the pitch, increase the sound intensity, and increase the speech rate, the acoustic feature prediction network is jointly trained using the normal speech and the training style embedding information, and the predicted features of the sound to be synthesized obtained may include increasing the volume, sound intensity, and speech rate by set values respectively based on the normal speech.
[0069] There can be various implementation manners for the acoustic feature prediction network. To improve the naturalness of the predicted features of the sound to be synthesized, Tacotron can be selected as the acoustic feature prediction network.
[0070] Please refer to Figure 3 , in one implementation, the acoustic feature prediction network may include: an encoder, a second attention module, and a decoder. Correspondingly, step S120 of inputting the text information to be output and the training style embedding information into the acoustic feature prediction network to obtain the predicted features of the sound to be synthesized corresponding to the home scene can be implemented in the following manner: input the text information to be output into the encoder to obtain fixed-length character embedding information; input the fixed-length character embedding information and the training style embedding information into the second attention module to obtain the aligned acoustic features and character information; input the aligned acoustic features and character information into the decoder to obtain the predicted features of the sound to be synthesized.
[0071] When Tacotron is selected as the acoustic feature prediction network, when the aligned acoustic features and character information are input into the decoder, the predicted features of the sound to be synthesized output by the decoder can be linear spectrum (Linear Spectrum) or mel spectrum (Mel Spectrum) or other acoustic features applicable to vocoders.
[0072] In this embodiment, the training style embedding information can be obtained through a scene acoustic style extractor. Please refer to Figure 4, the training style embedding information can be obtained through S210 and S220.
[0073] S210, obtain the acoustic features of the home scene in a noisy acoustic environment.
[0074] S210, input the acoustic features in the noisy acoustic environment into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene. Among them, the training style embedding information represents the scene acoustic style corresponding to the home scene.
[0075] The scene acoustic style in this embodiment may include one or more attributes. For example, it may include any one or any combination of two or more of the attributes such as pitch, sound intensity, speech rate, emotion, etc.
[0076] Among them, the home scene can be flexibly divided, and the acoustic features of the home scene in a noisy acoustic environment can be collected. For example, the kitchen, living room, and bathroom where noisy sounds often occur can be selected as the home scenes respectively and the acoustic features in the noisy acoustic environment can be collected. Another example is that the home exercise gym, multimedia audio-visual room, study, etc. can also be selected as the home scenes respectively and the acoustic features in the noisy acoustic environment can be collected.
[0077] It can be understood that the finer the division of the home scene, the more voices can be prepared according to each subdivided home scene, supporting the selection of more acoustic features for classification. Using the acoustic features in the noisy acoustic environment to train the scene acoustic style extractor, the obtained training style embedding information is more refined, and the scene acoustic style corresponding to each home scene represented is more delicate. Correspondingly, the scene style embedding information corresponding to the home scene determined by using the trained scene acoustic style extractor is more delicate, and the response speech data synthesized by the electronic device using the scene style embedding information is clearer and more natural.
[0078] The acoustic features of the home scene in a noisy acoustic environment are obtained based on the Lombard speech in the noisy home scene. For example, the Lombard speech in the noisy home scene can be collected and acoustic feature extraction can be performed, such as using an acoustic feature extraction module to perform acoustic feature extraction to obtain the acoustic features of the home scene in a noisy acoustic environment.
[0079] Lombard speech in a noisy home scenario can be obtained in various ways. Exemplarily, the speech of a user can be recorded under the interference of the ambient sound of various home scenarios as the training data for the scene acoustic style extractor. For example, volunteers can wear movable headphones, and the headphones play the ambient sound of different home scenarios. The ambient noise induces the volunteers to change their voices, and thus Lombard speech in different home scenarios can be captured as the training data for the scene acoustic style extractor.
[0080] To improve the training effect, the acoustic features obtained in the noisy home scenario can be as rich as possible. For example, by increasing the number of volunteers, the amount of speech of each volunteer in each noisy home scenario, etc., the Lombard speech captured in each noisy home scenario can be increased, thereby improving the richness of the acoustic features obtained in the noisy home scenario.
[0081] By obtaining the Lombard speech corresponding to the home scenario as richly and comprehensively as possible, extracting acoustic features, obtaining the acoustic features of the home scenario in a noisy acoustic environment, and inputting them into the scene acoustic style extractor for training, the accuracy of the training style embedding information extracted by the scene acoustic style extractor can be effectively improved. Correspondingly, the scene style embedding information corresponding to the home scenario determined by using the trained scene acoustic style extractor is more accurate and reliable.
[0082] In this embodiment, the acoustic features in a noisy acoustic environment can be various. For example, the acoustic features can be spectral envelope, fundamental frequency (F0) of the sound, acoustic energy, spectrogram such as Linear spectrogram, Mel Frequency Cepstrum Coefficient (MFCC), Mel spectrogram, etc.
[0083] In one implementation, the training style embedding information can be based only on Lombard speech. For example, for each home scenario, collect Lombard speech in the noisy home scenario and extract acoustic features to obtain the acoustic features of the home scenario in a noisy acoustic environment. Input the acoustic features in the noisy acoustic environment into the scene acoustic style extractor to extract the embedded representation (training style embedding information) of the corresponding scene acoustic style features.
[0084] The scene acoustic style extractor can have various implementation manners. Exemplarily, please refer to Figure 5 and Figure 6, the scene acoustic style extractor may include: a reference encoder, a first attention module, and a fully connected layer. Among them, for S220, the step of inputting the acoustic features in a noisy acoustic environment into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene can be implemented through S221, S222, and S223.
[0085] S221, input the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information.
[0086] S222, input the reference embedding information into the first attention module to obtain attention weights.
[0087] S223, input the attention weights into the fully connected layer to obtain the training style embedding information.
[0088] Based on this method, by inputting as rich and comprehensive acoustic features of the home scene in a noisy acoustic environment as possible into the scene acoustic style extractor, the training style embedding information can be obtained, and the training of the scene acoustic style extractor can be realized.
[0089] In another implementation, the training style embedding information can be obtained based on Lombard speech and the scene type information output by the first scene classification model.
[0090] Exemplarily, in the case of inputting the acoustic features of the home scene in a noisy acoustic environment into the scene acoustic style extractor, the environmental acoustic features of the home scene can also be obtained, input the environmental acoustic features of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene, and input the scene type information into the scene acoustic style extractor for fusion, that is, introduce the scene type information for the learning of the scene acoustic style extractor.
[0091] To improve the convenience and efficiency of fusion, the scene type information can be input into the scene acoustic style extractor in the form of weights for fusion.
[0092] For example, in the case where the first scene classification model is a VGG16 network, the step of inputting the environmental acoustic features of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene may include: inputting the environmental acoustic features of the home scene into the VGG16 network to obtain Softmax probability values; determining the first scene type weights corresponding to the Softmax probability values; using the first scene type weights as the scene type information.
[0093] Among them, the corresponding relationship between the Softmax probability value and the weight of the first scene type can be preset. Correspondingly, the step of determining the weight of the first scene type corresponding to the Softmax probability value can include: determining the weight of the first scene type corresponding to the Softmax probability value according to the corresponding relationship between the Softmax probability value and the weight of the first scene type.
[0094] Of course, in the case where the scene classification model is a VGG16 network, the label value can also be obtained.
[0095] Please refer to Figure 7 and Figure 8 In the case where the scene acoustic style extractor includes a reference encoder, a first attention module, and a fully connected layer, Figure 4 The step S220 shown in
[0096] i.e., the step of inputting the acoustic features in the noisy acoustic environment into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene can be implemented through S224, S225, and S226.
[0097] S224, input the acoustic features of the home scene in the noisy acoustic environment into the reference encoder to obtain the reference embedding information.
[0098] In the case of using the weight of the first scene type as the scene type information, correspondingly, the step of inputting the scene type information into the scene acoustic style extractor for fusion is implemented through S226.
[0099] S226, input the attention weight and the weight of the first scene type into the fully connected layer for weighting to obtain the training style embedding information.
[0100] For another example, in the case where the first scene classification model is a ResNet network, the step of inputting the environmental acoustic features of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene can include: inputting the environmental acoustic features of the home scene into the ResNet network to obtain the label value of the home scene; determining the weight of the second scene type corresponding to the label value; using the weight of the second scene type as the scene type information.
[0101] Among them, the corresponding relationship between the label value and the weight of the second scene type can be preset. Correspondingly, the step of determining the weight of the second scene type corresponding to the label value can include: determining the weight of the second scene type corresponding to the label value according to the corresponding relationship between the label value and the weight of the second scene type.
[0102] Please refer to Figure 9 and Figure 10, when the scene acoustic style extractor includes a reference encoder, a first attention module, and a fully connected layer, S220, the step of inputting the acoustic features in a noisy acoustic environment into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene can be implemented through S227, S228, and S229.
[0103] S227, input the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information.
[0104] S228, input the reference embedding information into the first attention module to obtain attention weights.
[0105] When the second scene type weight is used as the scene type information, correspondingly, the step of inputting the scene type information into the scene acoustic style extractor for fusion is implemented through S229.
[0106] S229, input the attention weights and the second scene type weights into the fully connected layer for weighting to obtain the training style embedding information.
[0107] The above process for obtaining the training style embedding information and the implementation structure of the scene acoustic style extractor are only examples. The process for obtaining the training style embedding information and the implementation structure of the scene acoustic style extractor can also be other, and this embodiment will not give examples one by one here. Based on the above process, the training of the scene acoustic style extractor can be realized.
[0108] In the above example description of the process for obtaining the training style embedding information, the step of inputting the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information may include: inputting the acoustic features of variable-length Lombardspeech (the acoustic features of the home scene in a noisy acoustic environment) into the reference encoder, and the reference encoder compresses the acoustic features of variable-length Lombard speech into fixed-length reference embedding information such as a vector of a fixed size. This reference embedding information is used to encode the overall acoustic style of an audio segment.
[0109] The reference encoder can have various implementation structures, and this embodiment will give the following examples.
[0110] For example, as Figure 11 shown, the reference encoder may include a convolutional neural network (Convolutional Neural Networks, CNN), a bidirectional long short-term memory recurrent neural network (including: a backward recurrent neural network (backward LSTM) and a forward recurrent neural network (forward LSTM)), and a mapping layer.
[0111] Input the acoustic features of the home scene in a noisy acoustic environment into a CNN to extract further acoustic features. Input the further acoustic features into a bidirectional long short-term memory recurrent neural network to obtain context-related acoustic features. Input the context-related acoustic features into a mapping layer to output fixed-length reference embedding information, encoding the overall acoustic style of an audio segment. Such as the Lombard speech acoustic style of the user in each home scene.
[0112] It can be understood that if different home scenes with similar environmental acoustic features are distinguished, correspondingly, the reference encoder processes the acoustic features of different home scenes with similar environmental acoustic features in a noisy acoustic environment, and can achieve refined processing of the acoustic styles corresponding to different home scenes with similar environmental acoustic features.
[0113] Another example, such as Figure 12 As shown, the reference encoder can include a first pre-trained model output layer, a CNN, and a mapping layer. The first pre-trained model can refer to the implementation of Audio word2vec, which is mainly implemented by a sequence-to-sequence autoencoder. The main structure in Audio word2vec is the Recurrent Neural Network (RNN). The first pre-trained model output layer outputs pre-trained vectors. The purpose of using Audio word2vec is to obtain a better expression of acoustic features. Then, through two to three convolutional neural networks, further acoustic features are extracted, and then two to three fully connected networks are used as the mapping layer to map to a predefined dimension.
[0114] Another example, such as Figure 13 As shown, the reference encoder can include a second pre-trained model output layer, a bidirectional long short-term memory recurrent neural network, and a mapping layer. The second pre-trained model can refer to the implementation of an unsupervised pre-trained model (wav2vec), which is mainly implemented by a CNN-based encoder. The second pre-trained model output layer outputs pre-trained vectors. The purpose of using Wave2Vec is to obtain a better expression of acoustic features. Then, through the bidirectional long short-term memory recurrent neural network, further context-related acoustic features are extracted, and then two to three fully connected networks are used as the mapping layer to map to a predefined dimension.
[0115] Using the first pre-trained model and the second pre-trained model can obtain more expressions about sound features.
[0116] The first attention module can obtain attention weights in various ways. Exemplarily, an acoustic style marker set can be formed by permuting and combining attributes such as pitch, sound intensity, speech rate, emotion, etc. During the process of the first attention module obtaining attention weights, a group of acoustic style markers is randomly selected from the acoustic style marker set for embedding. Each acoustic style marker can include one or more of the attributes such as pitch, sound intensity, speech rate, emotion, etc. The number of embedded acoustic style markers can be flexibly set. For example, it can be set to one, three, five, ten, etc., to represent a small number of different acoustic dimensions in the training data of the scene acoustic style extractor, such as one or more of the attributes like pitch, sound intensity, speech rate, emotion, etc. Using the reference embedding information as the query information of the first attention module, the first attention module learns the similarity metric between the reference embedding information and each acoustic style marker in the embedded group of acoustic style markers. The first attention module then outputs a group of combined weights (attention weights), and these combined weights represent the contribution of each acoustic style marker to the reference embedding information. Correspondingly, the training style embedding information output by the scene acoustic style extractor may be information related to any one or more of the attributes such as pitch, sound intensity, speech rate, emotion, etc. For example, in the case where a certain home scene is the living room, the training style embedding information corresponding to the living room may include increasing the pitch, increasing the sound intensity, and increasing the speech rate.
[0117] In this embodiment, the training style embedding information can have various presentation forms. For example, the training style embedding information can be specific values assigned to attributes such as pitch, sound intensity, speech rate, emotion, etc. For another example, the training style embedding information can be the adjustment values of attributes such as pitch, sound intensity, speech rate, emotion, etc. based on a certain normal speech. Among them, the normal speech can be speech uttered in a quiet scene.
[0118] Please refer to Figure 14 , in the case of introducing scene type information such as Softmax probability values or label values to obtain training style embedding information and training the scene acoustic style extractor, the Softmax probability values and label values can be processed into information that can be recognized and used by the fully connected layer through a style embedding regulator. For example, the first scene type weight corresponding to the Softmax probability value is determined through the style embedding regulator, and the second scene type weight corresponding to the label value is determined. In this case, the home scene to which the acoustic features in the noisy acoustic environment belong is known.
[0119] When the scene type information is a tag value, the fully connected layer can multiply the second scene type weight corresponding to the tag value by the attention weight, and only retain one acoustic style marker as the training style embedding information corresponding to the corresponding home scene, so as to judge which attribute has a greater influence on a certain acoustic style marker in the home scene. During training, manual definition of multiple groups of weight parameters is supported and directly called during inference, thereby improving the training efficiency.
[0120] When the scene type information is a Softmax probability value, the fully connected layer can multiply the first scene type weight corresponding to the softmax probability value by the attention weight to balance the contribution of each embedded acoustic style marker to the reference embedding information again, so that the scene acoustic style extractor can not only learn the acoustic style of Lombard speech in the corresponding home scene, but also learn the influence of different home scenes on the acoustic style.
[0121] In order to improve the acquisition efficiency of the training style embedding information and accelerate convergence, in another implementation, the second scene classification model can also be called. The second scene classification model can indicate the home scene type to which the acoustic features in a noisy acoustic environment belong. By inputting the acoustic features of the home scene in the noisy acoustic environment into the second scene classification model, the scene type indication information corresponding to the home scene is obtained, so that the training feedback of the fully connected layer can be carried out according to the scene type indication information. For example, during the process of inputting the acoustic features in the noisy acoustic environment into the scene acoustic style extractor to train the scene acoustic style extractor, the acoustic features in the noisy acoustic environment are input into the second scene classification model. The second scene classification model outputs the scene type indication information, and this scene type indication information is transmitted to the fully connected layer of the scene acoustic style extractor to perform training feedback on the fully connected layer, thereby assisting convergence and improving the acquisition efficiency of the training style embedding information. In this embodiment, when the second scene classification model is called to perform training feedback on the scene acoustic style extractor, the acoustic features in the noisy acoustic environment input into the second scene classification model and the acoustic features in the noisy acoustic environment input into the scene acoustic style extractor can be the same or different, as long as the acoustic features of the home scene input into the scene acoustic style extractor and the acoustic features of the home scene input into the second scene classification model in the noisy acoustic environment correspond to the same home scene.
[0122] The first scene classification model and the second scene classification model can be trained in various ways. Please refer to Figure 15 , which is a schematic flowchart of a method for training the first scene classification model provided by an embodiment of the present invention, and can be Figure 1The electronic device 100 shown performs, for example, and can be performed by the processor 120 in the electronic device 100. This first scenario classification model training method includes S310 and S320.
[0123] S310, obtain the environmental acoustic features of multiple home scenarios.
[0124] S320, input the environmental acoustic features of multiple home scenarios into a scenario classification network for training to obtain scenario type information corresponding to each home scenario.
[0125] Among them, home scenarios can be flexibly divided, and the environmental acoustic features of home scenarios can be collected. For example, the kitchen, living room, and bathroom where noisy sounds often occur can be selected as home scenarios respectively and the environmental acoustic features of the home scenarios can be collected. Another example is that the home exercise gym, multimedia audio-visual room, study, etc. can also be selected as home scenarios respectively and the environmental acoustic features of the home scenarios can be collected. By finely dividing the home scenarios and using the environmental acoustic features of different home scenarios to train the scenario classification network, scenario type information corresponding to each home scenario can be obtained, and the judgment of the home scenario type can be realized based on the scenario type information.
[0126] To improve the training effect, the environmental acoustic features of each home scenario obtained can be as rich as possible. For example, the environmental acoustic features of each home scenario can include the individual environmental acoustic features of each home device located in this home scenario, and the environmental acoustic features of at least pairwise permutations and combinations of each home device located in this home scenario. Exemplarily, when the home scenario is the kitchen, the environmental acoustic features of the kitchen can include the individual environmental acoustic features of each home device in the kitchen, such as the range hood, dishwasher, gas stove, pots and pans, etc., and the environmental acoustic features of the pairwise, three-way, four-way, etc. permutations and combinations of the range hood, dishwasher, gas stove, pots and pans, etc. in the kitchen.
[0127] By obtaining the environmental acoustic features of home scenarios as richly and comprehensively as possible and inputting them into the scenario classification network for training, the sensitivity and reliability of the trained first scenario classification model can be effectively improved. For example, by inputting the individual environmental acoustic features of home devices into the scenario classification network for training, the type of home scenario can be directly determined based on the environmental acoustic features of certain specific home devices subsequently. Exemplarily, based on the acoustic feature of the toilet flushing sound, the home scenario can be directly determined to be the bathroom. Another example is that by inputting the environmental acoustic features of various permutations and combinations of each home device in each home scenario into the scenario classification network for training, reliable discrimination between different home scenarios with similar environmental acoustic features can be achieved. Exemplarily, both the kitchen and the bathroom may have the acoustic feature of running water mixed in. However, combining the acoustic feature of the sound of stir-frying in a wok can determine that the home scenario is the kitchen.
[0128] Please refer to Figure 16 , in one implementation, S310, obtaining the environmental acoustic features of multiple home scenarios can be achieved through S311 and S312.
[0129] S311, obtaining the environmental acoustic data of multiple home scenarios.
[0130] S312, performing acoustic feature extraction on the environmental acoustic data to obtain the environmental acoustic features of multiple home scenarios.
[0131] Among them, acoustic feature extraction can be performed in various ways. For example, an acoustic feature extraction module can be used to perform acoustic feature extraction on the environmental acoustic data. The environmental acoustic data in this embodiment is the acoustic data generated during the use of each home device in the home scenario and does not include human voices.
[0132] The scene classification network can be flexibly selected. For example, it can be a set deep learning network. Exemplarily, the scene classification network can be a deep convolutional neural network, such as the VGG16 network, the ResNet network, etc. According to the different scene classification networks, different scene type information can be obtained. Exemplarily, when the scene classification network is the VGG16 network, in S320, the environmental acoustic features are input into the VGG16 network, and the obtained scene type information is the Softmax probability value (of course, a label value such as one-hot encoding can also be obtained). When the scene classification network is the ResNet network, in S320, the environmental acoustic features are input into the ResNet network, and the obtained scene type information is a label value such as one-hot encoding (of course, the obtained scene type information can also be the softmax probability value).
[0133] Based on the above design, according to the label value representing the home scene type, the specific home scene type to which the environmental acoustic features of the home scene belong can be determined, or according to the softmax probability value representing the home scene belonging to each type, different weight combinations of the home scene style can be determined, thereby realizing the refinement of the home scene recognition granularity, improving the accuracy of home scene recognition, and better "guiding" the extraction of scene style embedding.
[0134] Based on the refined classification of the home scene by the first scene classification model, a more detailed division of the scene acoustic style in the home scene can be made. And it can also better "guide" the extraction of scene style embedding. Exemplarily, when the home scene includes three types: kitchen, living room, and bathroom, the scene acoustic style extractor can perform scene acoustic style extraction for the kitchen, living room, and bathroom respectively in combination with the first scene classification model.
[0135] To more clearly illustrate the first scenario classification model training method in the embodiments of the present invention, the following scenario is used as an example for illustration.
[0136] In the case where the home scenario includes three types: kitchen, living room, and bathroom, collect the environmental acoustic data separately generated by each home appliance in the kitchen, as well as the environmental acoustic data generated by two permutations, three permutations, four permutations, etc., and use all the collected environmental acoustic data in the kitchen as the first acoustic data set. Collect the environmental acoustic data separately generated by each home appliance in the living room, as well as the environmental acoustic data generated by two permutations, three permutations, four permutations, etc., and use all the collected environmental acoustic data in the living room as the second acoustic data set. Collect the environmental acoustic data separately generated by each home appliance in the bathroom, as well as the environmental acoustic data generated by two permutations, three permutations, four permutations, etc., and use all the collected environmental acoustic data in the bathroom as the third acoustic data set.
[0137] Extract the acoustic features of the environmental acoustic data in the first acoustic data set to obtain the first acoustic feature set. Extract the acoustic features of the environmental acoustic data in the second acoustic data set to obtain the second acoustic feature set. Extract the acoustic features of the environmental acoustic data in the third acoustic data set to obtain the third acoustic feature set.
[0138] Set the scenario type information corresponding to the kitchen as the first type, the scenario type information corresponding to the living room as the second type, and the scenario type information corresponding to the bathroom as the third type.
[0139] When the acoustic features in the first acoustic feature set are input into the scenario classification network, use the first type as the target output of the scenario classification network. When the acoustic features in the second acoustic feature set are input into the scenario classification network, use the second type as the target output of the scenario classification network. When the acoustic features in the third acoustic feature set are input into the scenario classification network, use the third type as the target output of the scenario classification network. Based on this, train the scenario classification network until the convergence condition is reached, and then the required first scenario classification model can be obtained.
[0140] The convergence condition can be set flexibly. For example, it can be that the accuracy rate of obtaining the scenario type information corresponding to each home scenario reaches a preset value. Exemplarily, the acoustic features in the first acoustic feature set, the second acoustic feature set, and the third acoustic feature set can be divided into training data and test data, and the scenario classification network is trained based on the training data until the accuracy rate of obtaining the target output after inputting the test data reaches the preset value, and it is determined that the convergence condition is reached, and the required first scenario classification model is obtained. Another example is that the number of training times reaches a set amount. This embodiment does not limit this.
[0141] After obtaining the first scene classification model by using the above scene classification model training method, based on the first scene classification model, the weight combination of the specific home scene or home scene style to which the environmental acoustic features belong can be identified. For example, if the environmental acoustic features of a certain home scene are input into the scene classification model and the scene classification model outputs the first type, it can be determined that the environmental acoustic features belong to the kitchen according to the first type.
[0142] After inputting the environmental acoustic features of multiple home scenes into the trained first scene classification model, the first scene classification model outputs scene type-related information, such as label values or softmax probability values. The scene type information output by the first scene classification model can be used as the input of the scene acoustic style extractor for combined training and use of the scene acoustic style extractor.
[0143] The embodiment of the present invention also provides a second scene classification model training method, including: obtaining the acoustic features of multiple home scenes in a noisy acoustic environment, inputting the acoustic features of multiple home scenes in a noisy acoustic environment into a scene classification network for training, and obtaining scene type indication information corresponding to each home scene.
[0144] The second scene classification model uses the acoustic features of multiple home scenes in a noisy acoustic environment as input. Correspondingly, the second scene classification model can indicate the home scene type to which the acoustic features in the noisy acoustic environment belong.
[0145] Since the difference between the training processes of the second scene classification model and the first scene classification model lies only in the different input data, due to the different input data, there are differences in the output data. The scene type indication information output by the second scene classification model is used as the training feedback of the scene acoustic style extractor. For example, the fully connected layer of the scene acoustic style extractor is trained and feedback according to the scene type indication information to accelerate convergence. The scene type information output by the first scene classification model is used for weighting. For example, the scene type information and the attention weight are input into the fully connected layer of the scene acoustic style extractor for weighting. The optional structure and training principle of the second scene classification model can refer to the relevant description of the first scene classification model above, so it will not be elaborated here.
[0146] During the training and application process of the scene acoustic style extractor, the first scene classification model can be selected to be called, or the second scene classification model can be selected to be called, or both the first scene classification model and the second scene classification model can be selected to be called, or neither the first scene classification model nor the second scene classification model can be called.
[0147] Through the above solutions, the training of the acoustic feature prediction model, the scene acoustic style extractor, the acoustic feature prediction model, the first scene classification model, and the second scene classification model is achieved. Based on the above training process, the trained acoustic feature prediction model, scene acoustic style extractor, first scene classification model, and second scene classification model can be obtained. The above training methods for the first scene classification model, the second scene classification model, the scene acoustic style extractor, and the acoustic feature prediction model can be executed in the same electronic device or in different electronic devices. The training methods for the first scene classification model, the second scene classification model, the scene acoustic style extractor, and the acoustic feature prediction model can run independently or can be combined and run in different combinations according to requirements. The trained scene acoustic style extractor, acoustic feature prediction model, first scene classification model, and second scene classification model can run independently or can be combined and run in different combinations according to requirements.
[0148] Please refer to Figure 17 , which is a schematic flowchart of a speech synthesis method provided by an embodiment of the present invention and can be executed by Figure 1 the electronic device 100 shown, for example, can be executed by the processor 120 in the electronic device 100. The speech synthesis method includes S410, S420, and S430.
[0149] S410, determining the scene style embedding information corresponding to the home scene; wherein, the scene style embedding information represents the scene acoustic style corresponding to the home scene.
[0150] S420, determining the predicted acoustic features corresponding to the home scene according to the response text content information and the scene style embedding information.
[0151] S430, synthesizing the predicted acoustic features to obtain the output response speech data.
[0152] In one implementation, the electronic device can be a smart device in a smart home scenario. The output response speech data synthesized by the smart device through S410 to S430 is Lombard speech with the Lombard effect. The smart device outputs the output response speech data in a noisy home scenario so that it can be clearly transmitted to the user, ensuring smooth interaction. In another implementation, the electronic device can be a server, which can be communicatively connected to the smart device in the smart home scenario. The output response speech data synthesized by the server through S410 to S430 is Lombard speech with the Lombard effect. The server sends the synthesized output response speech data to the smart device, and the smart device plays the output response speech data in a noisy home scenario so that it can be clearly transmitted to the user, ensuring smooth interaction.
[0153] There can be various implementation manners for S410 to S430. For example, the scene style embedding information corresponding to the home scenario can be determined by a trained scene acoustic style extractor. For another example, the predicted acoustic features corresponding to the home scenario can be determined by a trained acoustic feature prediction model.
[0154] Please refer to Figure 18 , in one implementation, the electronic device can call a trained scene acoustic style extractor to obtain the acoustic features of the home scenario in a noisy acoustic environment, input the acoustic features of the home scenario in a noisy acoustic environment into the scene acoustic style extractor, and obtain the scene style embedding information corresponding to the home scenario. The electronic device can call a trained acoustic feature prediction model, use the response text content information and the scene style embedding information as the input of the acoustic feature prediction model, obtain the predicted acoustic features corresponding to the home scenario, and further synthesize the output response speech data.
[0155] Please refer to Figure 19 , in another implementation, the electronic device can call a trained scene acoustic style extractor and a first scene classification model, input the environmental acoustic features of the home scenario into the first scene classification model to obtain the scene type information corresponding to the home scenario, input the scene type information into the scene acoustic style extractor for fusion, such as using the scene type information and the acoustic features of the home scenario in a noisy acoustic environment as the input of the scene acoustic style extractor together, and the scene acoustic style extractor outputs the scene style embedding information corresponding to the home scenario. The electronic device can call a trained acoustic feature prediction model, use the response text content information and the scene style embedding information as the input of the acoustic feature prediction model, obtain the predicted acoustic features corresponding to the home scenario, and further synthesize the output response speech data.
[0156] For example, during speech synthesis, a trained scene acoustic style extractor and an acoustic feature prediction model can be invoked. Please refer to Figure 20 , the scene acoustic style extractor may include: a reference encoder, a first attention module, and a fully connected layer. In the case where only the acoustic features of the home scene in a noisy acoustic environment are used as the input to the scene acoustic style extractor, the scene style embedding information corresponding to the home scene can be obtained in the following way: input the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information; input the reference embedding information into the first attention module to obtain attention weights; input the attention weights into the fully connected layer to obtain the scene style embedding information. Use the response text content information and the scene style embedding information as the input to the acoustic feature prediction model, thereby obtaining the predicted acoustic features corresponding to the home scene, synthesizing the predicted acoustic features, and further obtaining the output response speech data.
[0157] For another example, during speech synthesis, a trained scene acoustic style extractor, a first scene classification model, and an acoustic feature prediction model can be invoked. Please refer to Figure 21 , the scene acoustic style extractor may include: a reference encoder, a first attention module, and a fully connected layer. In the case where the scene type information and the acoustic features of the home scene in a noisy acoustic environment are jointly used as the input to the scene acoustic style extractor, correspondingly, the scene style embedding information corresponding to the home scene can be obtained in the following way: input the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information; input the reference embedding information into the first attention module to obtain attention weights; input the attention weights and the scene type information into the fully connected layer for weighting to obtain the scene style embedding information. Use the response text content information and the scene style embedding information as the input to the acoustic feature prediction model, thereby obtaining the predicted acoustic features corresponding to the home scene, synthesizing the predicted acoustic features, and further obtaining the output response speech data.
[0158] When the first scene classification model is a VGG16 network, the environmental acoustic features of the home scene can be input into the VGG16 network to obtain a Softmax probability value; determine the first scene type weight corresponding to the Softmax probability value; use the first scene type weight as the scene type information. Among them, according to the corresponding relationship between the Softmax probability value and the first scene type weight, the first scene type weight corresponding to the Softmax probability value can be determined.
[0159] Correspondingly, by inputting the attention weights and the first scene type weights into the fully connected layer for weighting, the scene style embedding information can be obtained.
[0160] When the first scene classification model is a ResNet network, the environmental acoustic features of the home scene can be input into the ResNet network to obtain the label value of the home scene; determine the second scene type weight corresponding to the label value; and use the second scene type weight as the scene type information. Among them, the second scene type weight corresponding to the label value can be determined according to the corresponding relationship between the label value and the second scene type weight.
[0161] Correspondingly, by inputting the attention weight and the second scene type weight into the fully connected layer for weighting, the scene style embedding information can be obtained.
[0162] In this embodiment, the acoustic features of the home scene in a noisy acoustic environment are used as the input of the scene acoustic style extractor, or both the acoustic features of the home scene in a noisy acoustic environment and the scene type information output by the above first scene classification model are used as the input of the scene acoustic style extractor. The scene acoustic style extractor then outputs the scene style embedding information. The scene style embedding information output by the scene acoustic style extractor can be used as the input of the following acoustic feature prediction model.
[0163] Please refer to Figure 22 , the acoustic feature prediction model can include an encoder, a second attention module, and a decoder. Correspondingly, taking the response text content information and the scene style embedding information as the input of the acoustic feature prediction model, the predicted acoustic features corresponding to the home scene can be obtained in the following way: input the response text content information into the encoder to obtain the character embedding information with a fixed length; input the character embedding information with a fixed length and the scene style embedding information into the second attention module to obtain the aligned acoustic features and character information; input the aligned acoustic features and character information into the decoder to obtain the predicted acoustic features corresponding to the home scene.
[0164] In this embodiment, taking the response text content information and the scene style embedding information as the input of the acoustic feature prediction model, the acoustic feature prediction model then outputs the predicted acoustic features. Based on the predicted acoustic features, the output response speech data with the Lombard effect can be synthesized, so that the output response speech data played by the electronic device in a noisy home scene can be reliably received by the user, thus ensuring the smoothness of human-computer interaction in the smart home scene and better meeting the actual application requirements.
[0165] In one implementation, a speech synthesis device can be used to synthesize the output response speech data. Such as Figure 23As shown in the figure, a vocoder can be used as a speech synthesis device to synthesize the predicted acoustic features output by the acoustic feature prediction model. The predicted acoustic features are input into the vocoder to obtain output response speech data containing the information of the response text content. In this embodiment, the vocoder can be flexibly selected. For example, when the predicted acoustic feature is Linear Spectrum, Griffin-Lim can be selected as the vocoder for converting the spectrum to waveform, or WaveNet can be selected as the vocoder. Another example is that when the predicted acoustic feature is Mel Spectrum, WaveNet can be selected as the vocoder.
[0166] The above describes the optional embodiments of the speech synthesis method in the embodiments of the present invention. In other implementation manners, based on the same design concept: by "imitating" the communication mode in which humans actively change their vocalization methods under the Lombard effect, for different home scenarios, speech data with corresponding scene acoustic features is synthesized for broadcasting. By synthesizing Lombard speech with better naturalness and clarity, the smoothness of voice interaction with users in a noisy home environment is ensured. There may be other implementation manners for the speech synthesis method in the embodiments of the present invention.
[0167] In another implementation manner, please refer to Figure 24 , the Lombard speech can be input into the Lombard speech generation model, and the Lombard speech generation model is used to learn the Lombard speech, so that the Lombard speech generation model can directly obtain the output response speech data according to the information of the response text content. Based on this kind of Lombard speech generation model, the intelligent device does not need to train and call the first scene classification model, the second scene classification model, the scene acoustic style extractor, and the acoustic feature prediction model. The information of the response text content is directly input into the Lombard speech generation model, and the output response speech data corresponding to the home scene can be obtained.
[0168] In this embodiment, a speech synthesis device can be used to synthesize the predicted acoustic features to obtain the output response speech data. In one implementation manner, the speech synthesis device can be a vocoder, and the predicted acoustic features are input into the vocoder, and then the output response speech data is obtained.
[0169] In another implementation manner, the speech synthesis model can also be directly trained, and then the output response speech data is obtained based on the trained speech synthesis model.
[0170] Please refer to Figure 25 , which is a schematic flowchart of a speech synthesis model training method provided by an exemplary embodiment of the present invention. It can be Figure 1The electronic device 100 shown performs, for example, and can be performed by the processor 120 in the electronic device 100. The method for training the speech synthesis model includes S510, S520, and S530.
[0171] S510, obtain the text information to be output and the training style embedding information corresponding to the home scene.
[0172] S520, obtain the predicted sound features to be synthesized corresponding to the home scene according to the text information to be output and the training style embedding information;
[0173] S530, synthesize the predicted sound features to be synthesized to obtain the speech data to be output.
[0174] Among them, the training style embedding information corresponding to the home scene can be obtained through a scene acoustic style extractor. For example, the acoustic features in the noisy acoustic environment corresponding to the home scene can be input into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene. The predicted sound features to be synthesized can be obtained through an acoustic feature prediction network. For example, the training style embedding information and the text information to be output can be used as the input of the acoustic feature prediction network to obtain the predicted sound features to be synthesized corresponding to the home scene.
[0175] Correspondingly, please refer to Figure 26 , the acoustic features of the home scene in the noisy acoustic environment can be input into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene. Among them, the training style embedding information represents the scene acoustic style corresponding to the home scene. The training style embedding information and the text information to be output are used as the input of the acoustic feature prediction network to obtain the predicted sound features to be synthesized corresponding to the home scene. The predicted sound features to be synthesized are synthesized to obtain the speech data to be output.
[0176] In this embodiment, the scene acoustic style extractor may include: a reference encoder, a first attention module, and a fully connected layer. The training style embedding information can be obtained through the following steps: input the acoustic features of the home scene in the noisy acoustic environment into the reference encoder to obtain reference embedding information; input the reference embedding information into the first attention module to obtain attention weights; input the attention weights into the fully connected layer to obtain the training style embedding information. Please return to refer to Figure 5 and Figure 6 , one of the training processes of the scene acoustic style extractor can refer to the relevant descriptions of S210 and S221 to S223 above, and will not be elaborated here.
[0177] The acoustic feature prediction network may include an encoder, a second attention module, and a decoder. Correspondingly, the predicted features of the sound to be synthesized can be obtained through the following steps: inputting the text information to be output into the encoder to obtain character embedding information of a fixed length; inputting the character embedding information of the fixed length and the training style embedding information into the second attention module to obtain the aligned acoustic features and character information; inputting the aligned acoustic features and character information into the decoder to obtain the predicted features of the sound to be synthesized. Please refer back to Figure 2 and Figure 3 , the training process of the acoustic feature prediction model can refer to the relevant descriptions of S110 and S120 above, which will not be elaborated here.
[0178] On the basis of training the speech synthesis model by combining the training method of the scene acoustic style extractor and the training method of the acoustic feature prediction model, the above-mentioned first scene classification model training method can also be introduced to train the speech synthesis model. The training process of the first scene classification model can refer to the relevant descriptions of S310 and S320 above, which will not be elaborated here. Correspondingly, the speech synthesis model training method may further include: inputting the environmental acoustic features of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene, and inputting the scene type information into the scene acoustic style extractor for fusion.
[0179] In one implementation, when the first scene classification model is a VGG16 network, the step of inputting the environmental acoustic features of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene may include: inputting the environmental acoustic features of the home scene into the VGG16 network to obtain Softmax probability values; determining the first scene type weights corresponding to the Softmax probability values; using the first scene type weights as the scene type information. Exemplarily, the first scene type weights corresponding to the Softmax probability values can be determined according to the corresponding relationship between the Softmax probability values and the first scene type weights.
[0180] The scene acoustic style extractor may include: a reference encoder, a first attention module, and a fully connected layer. The training style embedding information can be obtained through the following method: inputting the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information; inputting the reference embedding information into the first attention module to obtain attention weights. Correspondingly, the step of inputting the scene type information into the scene acoustic style extractor for fusion includes: inputting the attention weights and the first scene type weights into the fully connected layer for weighting to obtain the training style embedding information. Please refer back to Figure 7 and Figure 8 , one of the training processes of the scene acoustic style extractor can refer to the relevant descriptions of S210 and S224 to S226 above, which will not be elaborated here.
[0181] In another implementation, when the scene classification model is a ResNet network, the steps of inputting the environmental acoustic features of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene may include: inputting the environmental acoustic features of the home scene into the ResNet network to obtain the label value of the home scene; determining the second scene type weight corresponding to the label value; and using the second scene type weight as the scene type information. Exemplarily, the second scene type weight corresponding to the label value may be determined according to the correspondence between the label value and the second scene type weight.
[0182] The scene acoustic style extractor may include: a reference encoder, a first attention module, and a fully connected layer. The training style embedding information may be obtained in the following manner: inputting the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain the reference embedding information; and inputting the reference embedding information into the first attention module to obtain the attention weight. Accordingly, the step of inputting the scene type information into the scene acoustic style extractor for fusion may include: inputting the attention weight and the second scene type weight into the fully connected layer for weighting to obtain the training style embedding information. Please refer back to Figure 9 and Figure 10 , one training process of the scene acoustic style extractor may refer to the relevant descriptions of S210 and S227 to S229 above, which will not be elaborated here.
[0183] In this embodiment, a second scene classification model training method may also be introduced. The acoustic features of the home scene in a noisy acoustic environment are input into the above-mentioned second scene classification model to obtain the scene type indication information corresponding to the home scene, and the fully connected layer is trained and fed back according to the scene type indication information to accelerate convergence.
[0184] A voice synthesis device may be used to synthesize the predicted features of the voice to be synthesized to obtain the voice data to be output. The voice synthesis device may be a vocoder. Accordingly, a vocoder may be used as the voice synthesis device to synthesize the predicted features of the voice to be synthesized output by the acoustic feature prediction network, and the predicted features of the voice to be synthesized are input into the vocoder to obtain the voice data to be output.
[0185] Based on the above, it can be seen that the speech synthesis model in this embodiment can be trained based on the aforementioned scene acoustic style extractor and acoustic feature prediction network, and the speech synthesis model can also be trained based on the aforementioned first scene classification model, second scene classification model, scene acoustic style extractor and acoustic feature prediction network. The training processes and implementation principles of the first scene classification model, second scene classification model, scene acoustic style extractor and acoustic feature prediction model applied in the speech synthesis model training method can be directly referred to the corresponding descriptions in the above scene classification model training method, scene acoustic style extractor training method and acoustic feature prediction model training method, and similar content will not be elaborated here.
[0186] Through the training of the speech synthesis model, after inputting the acoustic features and response text content information of the home scene in a noisy acoustic environment into the trained speech synthesis model, the speech synthesis model can output output response speech data containing the response text content information, and the output response speech data is Lombard speech with better naturalness and clarity.
[0187] To more clearly illustrate the speech synthesis method in the embodiments of the present invention, the following scenario is taken as an example for illustration.
[0188] The electronic device is a smart device. The smart device is located in a home scene and can perform voice interaction with the user, such as playing the output response speech data. The smart device is loaded with a first scene classification model, a second scene classification model, a scene acoustic style extractor and an acoustic feature prediction model.
[0189] The smart device determines whether the user has made a pre-configuration. If it is determined that the user has made a pre-configuration, the output response speech data is generated according to the pre-configuration; if it is determined that the user has not made a pre-configuration, the smart device receives the surrounding sounds and determines whether the received sound is only the ambient sound of the home scene or the user's speech in a noisy acoustic environment.
[0190] If it is determined that the received sound is only the ambient sound of the home scene, the first scene classification model is called, and the ambient sound of the home scene is input into the first scene classification model to obtain the scene type information corresponding to the home scene, such as the label value or softmax probability value, for classifying the home scene.
[0191] If it is determined that the received voice is the user's voice in a noisy acoustic environment, the intelligent device calls the scene acoustic style extractor and the second scene classification model. When the first scene classification model has not obtained the scene type information corresponding to the home scene, the intelligent device inputs the user's voice in the noisy acoustic environment into the scene acoustic style extractor and the second scene classification model to generate the corresponding scene style embedding information. When the first scene classification model has obtained the scene type information corresponding to the home scene, the intelligent device inputs the user's voice in the noisy acoustic environment and the scene type information corresponding to the home scene into the scene acoustic style extractor to generate the corresponding scene style embedding information.
[0192] Among them, before inputting the user's voice in the noisy acoustic environment into the scene acoustic style extractor, noise reduction processing can also be performed, and the denoised voice is input into the scene acoustic style extractor to further improve the accuracy of extracting the scene style embedding information.
[0193] When the intelligent device monitors that the user issues a question, it determines whether it has obtained the scene style embedding information corresponding to the home scene. If it is determined that it has obtained the scene style embedding information corresponding to the home scene, it calls the acoustic feature prediction model trained based on the to-be-output text information and the scene style embedding information, inputs the response text content information and the scene style embedding information corresponding to the home scene into the acoustic feature prediction model to generate the predicted acoustic features of Lombard speech in the corresponding home scene, and synthesizes and outputs the response voice data through a vocoder. If it is determined that it has not obtained the scene style embedding information corresponding to the home scene, it calls the Lombard speech generation model, inputs the response text content information into the Lombard speech generation model to generate the predicted acoustic features of Lombard speech in the corresponding home scene, and synthesizes and outputs the response voice data through a vocoder. The intelligent device broadcasts the synthesized output response voice data, so as to be able to interact with the user in a voice with the acoustic style in the corresponding home scene. The broadcast voice is Lombard speech with good naturalness and clarity, thus ensuring the smoothness of the voice interaction. It can be understood that the above voice interaction process can be repeatedly executed according to user needs.
[0194] To execute the corresponding steps in the above embodiments and various possible ways, an implementation manner of a voice synthesis device is given below. Please refer to Figure 27 , Figure 27 which is a functional module diagram of a first voice synthesis device 140 provided by an embodiment of the present invention. The first voice synthesis device 140 can be applied to Figure 1The electronic device 100 shown. It should be noted that the first voice synthesis device 140 provided in this embodiment has the same basic principle and technical effects as the above-mentioned voice synthesis method embodiment. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above-mentioned voice synthesis method embodiment. The first voice synthesis device 140 includes an information determination module 141 and a response voice data synthesis module 142.
[0195] Among them, the information determination module 141 is used to determine the scene style embedding information corresponding to the home scene, and determine the predicted acoustic features corresponding to the home scene according to the response text content information and the scene style embedding information; wherein, the scene style embedding information represents the scene acoustic style corresponding to the home scene.
[0196] The response voice data synthesis module 142 is used to synthesize the predicted acoustic features to obtain the output response voice data.
[0197] Please refer to Figure 28 , the implementation manner of the second voice synthesis device 150 is also given in the embodiment of the present invention. The second voice synthesis device 150 includes: a predicted acoustic feature obtaining module 151 and an information synthesis module 152.
[0198] Among them, the predicted acoustic feature obtaining module 151 is used to input the response text content information into the acoustic feature prediction model to obtain the predicted acoustic features corresponding to the home scene.
[0199] The information synthesis module 152 is used to synthesize the predicted acoustic features to obtain the output response voice data.
[0200] In order to execute the corresponding steps in the above embodiments and various possible manners, the implementation manner of a scene classification model training device is given below. Please refer to Figure 29 , Figure 29 is a functional module diagram of a scene classification model training device 160 provided in the embodiment of the present invention. The scene classification model training device 160 can be applied to Figure 1 the electronic device 100 shown. It should be noted that the scene classification model training device 160 provided in this embodiment has the same basic principle and technical effects as the above-mentioned scene classification model training method embodiment. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above-mentioned scene classification model training method embodiment. The scene classification model training device 160 includes an environmental acoustic feature obtaining module 161 and a scene classification network training module 162.
[0201] Among them, the environmental acoustic feature obtaining module 161 is used to obtain the environmental acoustic features of multiple home scenes.
[0202] The scene classification network training module 162 is used to input the environmental acoustic features of the multiple home scenes into the scene classification network for training, so as to obtain the scene type information corresponding to each home scene.
[0203] To execute the corresponding steps in the above embodiments and each possible way, an implementation manner of a scene acoustic style extractor training device is given below. Please refer to Figure 30 , Figure 30 FIG. is a functional module diagram of a scene acoustic style extractor training device 170 provided by an embodiment of the present invention. The scene acoustic style extractor training device 170 can be applied to Figure 1 the electronic device 100 shown in FIG.. It should be noted that for the scene acoustic style extractor training device 170 provided in this embodiment, its basic principle and the technical effects produced are the same as those in the above embodiments of the scene acoustic style extractor training method. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above embodiments of the scene acoustic style extractor training method. The scene acoustic style extractor training device 170 includes a noisy environment acoustic feature obtaining module 171 and a scene acoustic style extractor training module 172.
[0204] Among them, the noisy environment acoustic feature obtaining module 171 is used to obtain the acoustic features of the home scene in a noisy acoustic environment.
[0205] The scene acoustic style extractor training module 172 is used to input the acoustic features of the home scene in a noisy acoustic environment into the scene acoustic style extractor, so as to obtain the training style embedding information corresponding to the home scene; wherein, the training style embedding information represents the scene acoustic style corresponding to the home scene.
[0206] To execute the corresponding steps in the above embodiments and each possible way, an implementation manner of an acoustic feature prediction model training device is given below. Please refer to Figure 31 , Figure 31 FIG. is a functional module diagram of an acoustic feature prediction model training device 180 provided by an embodiment of the present invention. The acoustic feature prediction model training device 180 can be applied to Figure 1 the electronic device 100 shown in FIG.. It should be noted that for the acoustic feature prediction model training device 180 provided in this embodiment, its basic principle and the technical effects produced are the same as those in the above embodiments of the acoustic feature prediction model training method. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above embodiments of the acoustic feature prediction model training method. The acoustic feature prediction model training device 180 includes a data obtaining module 181 and an acoustic feature prediction model training module 182.
[0207] Among them, the data acquisition module 181 is used to acquire the text information to be output and the scene style embedding information corresponding to the home scene.
[0208] The acoustic feature prediction model training module 182 is used to input the text information to be output and the training style embedding information into the acoustic feature prediction network to obtain the predicted acoustic features of the sound to be synthesized corresponding to the home scene.
[0209] The acoustic feature prediction network includes an encoder, a second attention module, and a decoder; the acoustic feature prediction model training module 182 is used to input the text information to be output and the training style embedding information into the acoustic feature prediction network through the following steps to obtain the predicted acoustic features of the sound to be synthesized corresponding to the home scene:
[0210] Input the text information to be output into the encoder to obtain character embedding information of a fixed length;
[0211] Input the character embedding information of the fixed length and the training style embedding information into the second attention module to obtain the aligned acoustic features and character information;
[0212] Input the aligned acoustic features and character information into the decoder to obtain the predicted acoustic features of the sound to be synthesized corresponding to the home scene.
[0213] The data acquisition module 181 is used to obtain the training style embedding information corresponding to the home scene through the following steps:
[0214] Obtain the acoustic features of the home scene in a noisy acoustic environment;
[0215] Input the acoustic features of the home scene in a noisy acoustic environment into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene; among them, the training style embedding information represents the scene acoustic style corresponding to the home scene.
[0216] The data acquisition module 181 is also used for: inputting the environmental acoustic features of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene; inputting the scene type information into the scene acoustic style extractor for fusion.
[0217] When the first scene classification model is the VGG16 network, the data acquisition module 181 is used to input the environmental acoustic features of the home scene into the first scene classification model through the following steps to obtain the scene type information corresponding to the home scene:
[0218] Input the environmental acoustic features of the home scene into the VGG16 network to obtain the Softmax probability value;
[0219] Determine the first scene type weight corresponding to the Softmax probability value;
[0220] Use the first scene type weight as the scene type information.
[0221] The scene acoustic style extractor includes: a reference encoder, a first attention module, and a fully connected layer; the data acquisition module 181 is used to input the acoustic features of the home scene in a noisy acoustic environment into the scene acoustic style extractor through the following steps to obtain the training style embedding information corresponding to the home scene:
[0222] Input the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information;
[0223] Input the reference embedding information into the first attention module to obtain attention weights;
[0224] Input the attention weights and the first scene type weight into the fully connected layer for weighting to obtain the training style embedding information.
[0225] The data acquisition module 181 is used to determine the first scene type weight corresponding to the Softmax probability value through the following steps: Determine the first scene type weight corresponding to the Softmax probability value according to the correspondence between the Softmax probability value and the first scene type weight.
[0226] When the first scene classification model is a ResNet network, the data acquisition module 181 is used to input the environmental acoustic features of the home scene into the first scene classification model through the following steps to obtain the scene type information corresponding to the home scene:
[0227] Input the environmental acoustic features of the home scene into the ResNet network to obtain the label value of the home scene;
[0228] Determine the second scene type weight corresponding to the label value;
[0229] Use the second scene type weight as the scene type information.
[0230] The scene acoustic style extractor includes: a reference encoder, a first attention module, and a fully connected layer; the data acquisition module 181 is used to input the acoustic features of the home scene in a noisy acoustic environment into the scene acoustic style extractor through the following steps to obtain the training style embedding information corresponding to the home scene:
[0231] Input the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information;
[0232] Input the reference embedding information into the first attention module to obtain attention weights;
[0233] Input the attention weights and the second scene type weights into the fully connected layer for weighting to obtain the training style embedding information.
[0234] The data acquisition module 181 is used to determine the second scene type weights corresponding to the label values through the following steps: determine the second scene type weights corresponding to the label values according to the correspondence between the label values and the second scene type weights.
[0235] The data acquisition module 181 is also used to train the first scene classification model through the following steps: obtain the environmental acoustic features of multiple home scenes; input the environmental acoustic features of the multiple home scenes into a scene classification network for training to obtain the scene type information corresponding to each home scene.
[0236] The scene acoustic style extractor includes: a reference encoder, a first attention module, and a fully connected layer; the data acquisition module 181 is used to input the acoustic features of the home scene in a noisy acoustic environment into the scene acoustic style extractor through the following steps to obtain the training style embedding information corresponding to the home scene:
[0237] Input the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information;
[0238] Input the reference embedding information into the first attention module to obtain attention weights;
[0239] Input the attention weights into the fully connected layer to obtain the training style embedding information.
[0240] The data acquisition module 181 is also used to perform the following steps: input the acoustic features of the home scene in a noisy acoustic environment into the second scene classification model to obtain the scene type indication information corresponding to the home scene; perform training feedback on the fully connected layer according to the scene type indication information.
[0241] The data acquisition module 181 is also used to train the second scene classification model through the following steps: obtain the acoustic features of multiple home scenes in a noisy acoustic environment; input the acoustic features of the multiple home scenes in a noisy acoustic environment into a scene classification network for training to obtain the scene type indication information corresponding to each home scene.
[0242] On the basis described above, an embodiment of the present invention further provides a computer-readable storage medium, where the computer-readable storage medium includes a computer program, and when the computer program runs, it controls an electronic device where the computer-readable storage medium is located to execute the above-mentioned acoustic feature prediction model training method.
[0243] To execute the corresponding steps in the above embodiments and each possible manner, an implementation manner of a speech synthesis model training device is given below. Please refer to Figure 32 , Figure 32 FIG. is a functional module diagram of a speech synthesis model training device 190 provided by an embodiment of the present invention. The speech synthesis model training device 190 can be applied to Figure 1 the electronic device 100 shown. It should be noted that for the speech synthesis model training device 190 provided in this embodiment, its basic principle and the technical effects produced are the same as those in the above-mentioned speech synthesis model training method embodiment. For a brief description, for parts not mentioned in this embodiment, reference can be made to the corresponding content in the above-mentioned speech synthesis model training method embodiment. The speech synthesis model training device 190 includes an information acquisition module 191 and a to-be-output speech data synthesis module 192.
[0244] Among them, the information acquisition module 191 is configured to acquire to-be-output text information and scene style embedding information corresponding to a home scene, and obtain to-be-synthesized sound prediction features corresponding to the home scene according to the to-be-output text information and the scene style embedding information.
[0245] The to-be-output speech data synthesis module 192 is configured to synthesize the to-be-synthesized sound prediction features to obtain to-be-output speech data.
[0246] In the embodiment of the present invention, the synthesized output response speech data is Lombard speech with good naturalness and clarity, which can ensure the smoothness of speech interaction and improve the user experience.
[0247] In several embodiments provided by the present invention, it should be understood that the disclosed device and method can also be implemented in other ways. The device embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0248] In addition, each functional module in various embodiments of the present invention may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.
[0249] If the above-mentioned function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0250] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for training an acoustic feature prediction model, characterized in that, it includes: Obtain the text information to be output and the training style embedding information corresponding to the home scene; The training style embedding information is obtained by fusing the acoustic features of the home scene in a noisy acoustic environment and the scene type information corresponding to the home scene into a scene acoustic style extractor; Input the text information to be output and the training style embedding information into an acoustic feature prediction network to obtain the predicted features of the sound to be synthesized corresponding to the home scene; The step of obtaining the scene type information corresponding to the home scene includes: Obtain the environmental acoustic features of the home scene; Input the environmental acoustic features of the home scene into a first scene classification model to obtain the scene type information corresponding to the home scene.
2. The method for training an acoustic feature prediction model according to claim 1, characterized in that, The acoustic feature prediction network includes an encoder, a second attention module and a decoder; the step of inputting the text information to be output and the training style embedding information into the acoustic feature prediction network to obtain the predicted features of the sound to be synthesized corresponding to the home scene includes: Input the text information to be output into the encoder to obtain character embedding information of a fixed length; Input the character embedding information of the fixed length and the training style embedding information into the second attention module to obtain the aligned acoustic features and character information; Input the aligned acoustic features and character information into the decoder to obtain the predicted features of the sound to be synthesized corresponding to the home scene.
3. The method for training an acoustic feature prediction model according to claim 2, characterized in that, The training style embedding information corresponding to the home scene is obtained through the following steps: Obtain the acoustic features of the home scene in a noisy acoustic environment; Input the acoustic features of the home scene in the noisy acoustic environment into a scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene; wherein, the training style embedding information represents the scene acoustic style corresponding to the home scene.
4. The method for training an acoustic feature prediction model according to claim 1, characterized in that, When the first scene classification model is a VGG16 network, the scene type information is the first scene type weight corresponding to the Softmax probability value output by the VGG16 network.
5. The method for training an acoustic feature prediction model according to claim 1, characterized in that, When the first scene classification model is a ResNet network, the scene type information is the second scene type weight corresponding to the label value output by the ResNet network.
6. The method for training an acoustic feature prediction model according to claim 3, characterized in that, It further includes: Input the acoustic features of the home scene in a noisy acoustic environment into a second scene classification model to obtain the scene type indication information corresponding to the home scene; Perform training feedback on the scene acoustic style extractor according to the scene type indication information.
7. An apparatus for training an acoustic feature prediction model, characterized in that, it includes: A data acquisition module, configured to acquire the text information to be output and the training style embedding information corresponding to the home scene; The training style embedding information is obtained by inputting the acoustic features of the home scene in a noisy acoustic environment and the scene type information corresponding to the home scene into a scene acoustic style extractor for fusion; An acoustic feature prediction model training module, configured to input the text information to be output and the training style embedding information into an acoustic feature prediction network to obtain the predicted acoustic features of the sound to be synthesized corresponding to the home scene; The steps of obtaining the scene type information corresponding to the home scene include: Obtaining the environmental acoustic features of the home scene; Inputting the environmental acoustic features of the home scene into a first scene classification model to obtain the scene type information corresponding to the home scene.
8. An electronic device, characterized in that it includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the acoustic feature prediction model training method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that the computer-readable storage medium includes a computer program, and when the computer program runs, it controls the electronic device where the computer-readable storage medium is located to execute the acoustic feature prediction model training method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Speech synthesis model training method, speech synthesis method and device, equipment and storage medium
CN110264991A
Voice processing method and device, electronic equipment and storage medium
CN111326136A
Phrase-based end-to-end text-to-speech (TTS) synthesis
CN111681641A