Speech synthesis model training method, device, electronic device and computer-readable storage medium

By training the speech synthesis model to imitate the way humans speak under the Lombard effect and adjusting the speech features to adapt to noisy home environments, the problem of smart device voice broadcasts being drowned out by noise is solved, and clear voice broadcasts and smooth human-computer interaction are achieved in noisy environments.

CN114863908BActive Publication Date: 2025-06-03MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110069951.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-19
Publication Date
2025-06-03
Estimated Expiration
2041-01-19

AI Technical Summary

Technical Problem

In a noisy home environment, the voice broadcasts of smart devices are often drowned out by ambient noise, resulting in users being unable to receive them accurately and affecting the smoothness of human-computer interaction.

Method used

By imitating the way humans speak under the Lombard effect, the speech synthesis model is trained to actively adjust speech features in different acoustic environments and synthesize speech data with the acoustic style of the corresponding scene, ensuring speech clarity and naturalness.

Benefits of technology

In noisy environments, smart devices can broadcast voices clearly, improving the smoothness of human-computer interaction and ensuring that users can accurately receive voice information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114863908B_ABST
    Figure CN114863908B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a method, apparatus, electronic device, and computer-readable storage medium for training a speech synthesis model, which relates to the technical field of data processing. The method includes: obtaining text information to be output and training style embedding information corresponding to a home scene, obtaining a predicted sound feature to be synthesized corresponding to the home scene according to the text information to be output and the training style embedding information, and synthesizing the predicted sound feature to be synthesized to obtain speech data to be output, so as to synthesize speech that is adapted to the home scene and has good clarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a method, device, electronic device, and computer-readable storage medium for training a speech synthesis model. Background Art

[0002] Currently, electronic devices in a home environment, such as smart devices, can sense changes in the acoustic environment state and provide reasonable answers to users' questions. For example, when the smart device detects that the user issues an instruction message "turn on the microwave oven", it controls the microwave oven to turn on and announces "The microwave oven has been turned on" through voice, completing the response to the instruction message. However, through research, it is found that the content announced by the smart device cannot be accurately received by the user in many cases. For example, the announced voice is drowned out by the ambient sound in a noisy environment, making it impossible for the user to accurately receive the announced voice, which affects the smoothness of human-computer interaction in the smart home scenario and cannot meet the actual application requirements. Summary of the Invention

[0003] In 1909, the French otolaryngologist Étienne Lombard discovered through research that when communicating in a noisy environment, the speaker has to actively change the vocalization method to improve the sound effect and hopes that the other party can hear clearly. Through research, it is found that even if the same person pronounces the same voice, the voice characteristics are different in different environments. The changed characteristics include increasing the pitch, tone, loudness, and formant characteristics of the voice. This phenomenon is called the Lombard effect. In view of this, the inventor concludes that the prerequisite for the user to accurately receive reasonable content such as the voice announcement of "The microwave oven has been turned on" is that the smart device can change the vocalization method in different acoustic environments and actively improve the clarity and naturalness of the synthesized voice. Therefore, the inventor conducted research on this and further proposed a method for training a speech synthesis model that enables the smart device to "imitate" this change in the way humans actively change the vocalization method under the Lombard effect when the home environment type is noisy, and synthesize voice data with the acoustic style of the corresponding scenario for announcement, so as to ensure the smoothness of voice interaction with the user in a noisy home environment. Here, we refer to the synthesized voice with better recognition, naturalness, and intelligibility as Lombard speech.

[0004] One of the objectives of the present invention includes, for example, providing a method, device, electronic device, and computer-readable storage medium for training a speech synthesis model to at least partially improve the clarity of the synthesized voice.

[0005] Embodiments of the present invention may be implemented as follows:

[0006] In a first aspect, an embodiment of the present invention provides a method for training a speech synthesis model, including:

[0007] Obtaining the text information to be output and the training style embedding information corresponding to the home scene; wherein, the training style embedding information represents the scene acoustic style corresponding to the home scene;

[0008] Obtaining the predicted sound features to be synthesized corresponding to the home scene according to the text information to be output and the training style embedding information;

[0009] Synthesizing the predicted sound features to be synthesized to obtain the speech data to be output.

[0010] Based on two dimensions of the text information to be output and the training style embedding information, reliable prediction of the acoustic features corresponding to the home scene is realized, and the predicted sound features to be synthesized with the Lombard speech acoustic style are output. Furthermore, synthesizing the predicted sound features to be synthesized can improve the matching degree between the synthesized speech and the home scene, ensure the clarity of the speech broadcast by the intelligent device, and thus ensure that the broadcast speech can be accurately received by the user.

[0011] In a second aspect, an embodiment of the present invention provides a speech synthesis model training device, including:

[0012] An information acquisition module, configured to obtain the text information to be output and the training style embedding information corresponding to the home scene, and obtain the predicted sound features to be synthesized corresponding to the home scene according to the text information to be output and the training style embedding information; wherein, the training style embedding information represents the scene acoustic style corresponding to the home scene;

[0013] A speech data synthesis module to be output, configured to synthesize the predicted sound features to be synthesized to obtain the speech data to be output.

[0014] Based on two dimensions of the text information to be output and the training style embedding information, reliable prediction of the acoustic features corresponding to the home scene is realized, and the predicted sound features to be synthesized with the Lombard speech acoustic style are output. Furthermore, synthesizing the predicted sound features to be synthesized can improve the matching degree between the synthesized speech and the home scene, ensure the clarity of the speech broadcast by the intelligent device, and thus ensure that the broadcast speech can be accurately received by the user.

[0015] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the speech synthesis model training method according to any one of the foregoing embodiments. Correspondingly, this electronic device includes the beneficial effects in the speech synthesis model training method.

[0016] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a computer program. When the computer program runs, it controls an electronic device where the computer-readable storage medium is located to execute the voice synthesis model training method according to any one of the foregoing embodiments. Accordingly, the computer-readable storage medium includes the beneficial effects in the voice synthesis model training method.

[0017] To make the above objects, features, and advantages of the embodiments of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, provides a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 FIG. shows a schematic diagram of an application scenario provided by an embodiment of the present invention.

[0020] Figure 2 FIG. shows a schematic flowchart of a voice synthesis model training method provided by an embodiment of the present invention.

[0021] Figure 3 FIG. shows a schematic diagram of a training architecture of a voice synthesis model provided by an embodiment of the present invention.

[0022] Figure 4 FIG. shows a schematic flowchart of a training method for a scene acoustic style extractor provided by an embodiment of the present invention.

[0023] Figure 5 FIG. shows a schematic diagram of a training architecture of a scene acoustic style extractor provided by an embodiment of the present invention.

[0024] Figure 6 FIG. shows another schematic flowchart of a training method for a scene acoustic style extractor provided by an embodiment of the present invention.

[0025] Figure 7 FIG. shows another schematic diagram of a training architecture of a scene acoustic style extractor provided by an embodiment of the present invention.

[0026] Figure 8 FIG. shows another schematic flowchart of a training method for a scene acoustic style extractor provided by an embodiment of the present invention.

[0027] Figure 9Shows another schematic diagram of the training architecture of a scene acoustic style extractor provided by an embodiment of the present invention.

[0028] Figure 10 Shows another schematic flowchart of a method for training a scene acoustic style extractor provided by an embodiment of the present invention.

[0029] Figure 11 Shows one of the schematic diagrams of a reference encoder provided by an embodiment of the present invention.

[0030] Figure 12 Shows another schematic diagram of a reference encoder provided by an embodiment of the present invention.

[0031] Figure 13 Shows a third schematic diagram of a reference encoder provided by an embodiment of the present invention.

[0032] Figure 14 Shows another schematic diagram of the training architecture of a scene acoustic style extractor provided by an embodiment of the present invention.

[0033] Figure 15 Shows a schematic flowchart of a method for training a first scene classification model provided by an embodiment of the present invention.

[0034] Figure 16 Shows another schematic flowchart of a method for training a first scene classification model provided by an embodiment of the present invention.

[0035] Figure 17 Shows a schematic flowchart of a method for training an acoustic feature prediction model provided by an embodiment of the present invention.

[0036] Figure 18 Shows a schematic diagram of the training architecture of an acoustic feature prediction model provided by an embodiment of the present invention.

[0037] Figure 19 Shows a schematic flowchart of a method for speech synthesis provided by an embodiment of the present invention.

[0038] Figure 20 Shows a schematic diagram of the implementation principle of a method for speech synthesis provided by an embodiment of the present invention.

[0039] Figure 21 Shows another schematic diagram of the implementation principle of a method for speech synthesis provided by an embodiment of the present invention.

[0040] Figure 22 Shows a schematic diagram of the implementation principle of a scene acoustic style extractor provided by an embodiment of the present invention.

[0041] Figure 23Shows another schematic diagram of the implementation principle of a scene acoustic style extractor provided by an embodiment of the present invention.

[0042] Figure 24 Shows a schematic diagram of the implementation principle of an acoustic feature prediction model provided by an embodiment of the present invention.

[0043] Figure 25 Shows a schematic diagram of the implementation principle of a speech synthesis method provided by an embodiment of the present invention.

[0044] Figure 26 Shows a schematic diagram of the implementation principle of another method for synthesizing and outputting response speech data provided by an embodiment of the present invention.

[0045] Figure 27 Shows an exemplary structural block diagram of a first speech synthesis device provided by an embodiment of the present invention.

[0046] Figure 28 Shows an exemplary structural block diagram of a second speech synthesis device provided by an embodiment of the present invention.

[0047] Figure 29 Shows an exemplary structural block diagram of a scene classification model training device provided by an embodiment of the present invention.

[0048] Figure 30 Shows an exemplary structural block diagram of a scene acoustic style extractor training device provided by an embodiment of the present invention.

[0049] Figure 31 Shows an exemplary structural block diagram of an acoustic feature prediction model training device provided by an embodiment of the present invention.

[0050] Figure 32 Shows an exemplary structural block diagram of a speech synthesis model training device provided by an embodiment of the present invention.

[0051] Icons: Icon: 100 - Electronic device; 110 - Memory; 120 - Processor; 130 - Communication module; 140 - First speech synthesis device; 141 - Information determination module; 142 - Response speech data synthesis module; 150 - Second speech synthesis device; 151 - Predicted acoustic feature acquisition module; 152 - Information synthesis module; 160 - Scene classification model training device; 161 - Environmental acoustic feature acquisition module; 162 - Scene classification network training module; 170 - Scene acoustic style extractor training device; 171 - Noisy environment acoustic feature acquisition module; 172 - Scene acoustic style extractor training module; 180 - Acoustic feature prediction model training device; 181 - Data acquisition module; 182 - Acoustic feature prediction model training module; 190 - Speech synthesis model training device; 191 - Information acquisition module; 192 - To-be-output speech data synthesis module. Detailed implementation manners

[0052] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated herein usually can be arranged and designed in various different configurations.

[0053] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0054] It should be noted that the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0055] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it will not be repeatedly defined and explained in subsequent drawings without special instructions.

[0056] It should be noted that the features in the embodiments of the present invention can be combined with each other without conflict.

[0057] Please refer to Figure 1, which is a block diagram of an electronic device 100 provided in this embodiment. The electronic device 100 in this embodiment can be various processing devices capable of information interaction and processing. For example, the electronic device 100 can be an intelligent device capable of voice interaction with users. For another example, the electronic device can be a server, a processing platform, etc. capable of model training and speech synthesis.

[0058] The electronic device 100 may include a memory 110, a processor 120, and a communication module 130. The elements of the memory 110, the processor 120, and the communication module 130 are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines.

[0059] Among them, the memory 110 is used to store programs or data. The memory 110 can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.

[0060] The processor 120 is used to read / write the data or programs stored in the memory 110 and execute corresponding functions.

[0061] The communication module 130 is used to establish a communication connection between the electronic device 100 and other communication terminals through the network and is used to transmit and receive data through the network.

[0062] It should be understood that Figure 1 The structure shown is only a schematic diagram of the structure of the electronic device 100, and the electronic device 100 may also include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown. Figure 1 Each component shown can be implemented by hardware, software, or a combination thereof. For example, in the case where the electronic device 100 is an intelligent device capable of voice interaction with users, the electronic device 100 may further include a voice module, and the electronic device 100 can collect and output voices through the voice module.

[0063] In daily life, various noises fill the home environment, and people often need to communicate in noisy surroundings. Through research, it has been found that when communicating in a noisy environment, speakers will actively change the pitch or curvature of their voices to overcome the surrounding noise and improve speech clarity. This is the Lombard effect. With the rise of smart homes, how to improve the smoothness of human-machine interaction to better serve humans has become an issue of concern in this field.

[0064] In view of this, the inventor has studied how to improve the clarity of the speech broadcast by electronic devices such as smart devices in a noisy acoustic environment, and then proposed a method for training a speech synthesis model. The speech synthesis model is trained by "imitating" the communication mode in which humans actively change their vocalization methods under the Lombard effect. For different home scenarios, a speech synthesis model that can synthesize speech data with the acoustic characteristics of the corresponding scenario is trained. Then, through the trained speech synthesis model, Lombard speech with better naturalness and clarity can be synthesized, ensuring the smoothness of voice interaction with users in a noisy home environment.

[0065] Please refer to Figure 2 , which is a schematic flowchart of a method for training a speech synthesis model provided by an exemplary embodiment of the present invention, and can be executed by the electronic device 100 shown in Figure 1 . For example, it can be executed by the processor 120 in the electronic device 100. The method for training the speech synthesis model includes S110, S120, and S130.

[0066] S110, obtain the text information to be output and the training style embedding information corresponding to the home scenario.

[0067] S120, obtain the predicted sound features to be synthesized corresponding to the home scenario according to the text information to be output and the training style embedding information;

[0068] S130, synthesize the predicted sound features to be synthesized to obtain the speech data to be output.

[0069] There are various implementation manners for S110 to S130. For example, please refer to Figure 3 . The acoustic characteristics of the home scenario in a noisy acoustic environment can be input into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scenario. Among them, the training style embedding information represents the scene acoustic style corresponding to the home scenario. The training style embedding information and the text information to be output are used as the input of the acoustic feature prediction network to obtain the predicted sound features to be synthesized corresponding to the home scenario. The predicted sound features to be synthesized are synthesized to obtain the speech data to be output.

[0070] Please refer to Figure 4, the training style embedding information can be obtained through S210 and S220.

[0071] S210, obtain the acoustic features of the home scene in a noisy acoustic environment.

[0072] S210, input the acoustic features in the noisy acoustic environment into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene. Among them, the training style embedding information represents the scene acoustic style corresponding to the home scene.

[0073] The scene acoustic style in this embodiment may include one or more attributes. For example, it may include any one or a combination of two or more of the attributes such as pitch, sound intensity, speech rate, emotion, etc.

[0074] Among them, the home scene can be flexibly divided, and the acoustic features of the home scene in a noisy acoustic environment can be collected. For example, the kitchen, living room, and bathroom where noisy sounds often occur can be selected as the home scene respectively and the acoustic features in the noisy acoustic environment can be collected. Another example is that the home exercise gym, multimedia audio-visual room, study, etc. can also be selected as the home scene respectively and the acoustic features in the noisy acoustic environment can be collected.

[0075] It can be understood that the finer the division of the home scene, the more voices can be prepared according to each subdivided home scene, supporting the selection of more acoustic features for classification. Training the scene acoustic style extractor with the acoustic features in the noisy acoustic environment, the obtained training style embedding information is more refined, and the scene acoustic style corresponding to each home scene represented is more delicate. Correspondingly, the scene style embedding information corresponding to the home scene determined by using the trained scene acoustic style extractor is more delicate, and the response voice data synthesized by the electronic device using the scene style embedding information is clearer and more natural.

[0076] The acoustic features of the home scene in a noisy acoustic environment are based on Lombard speech in the noisy home scene. For example, Lombard speech in the noisy home scene can be collected and acoustic feature extraction can be performed, such as using an acoustic feature extraction module to perform acoustic feature extraction to obtain the acoustic features of the home scene in a noisy acoustic environment.

[0077] Lombard speech in a noisy home scene can be obtained in various ways. Exemplarily, the speech of a user can be recorded under the interference of the ambient sound of various home scenes as the training data of the scene acoustic style extractor. For example, volunteers can wear movable headphones, and the headphones play the ambient sound of different home scenes. The ambient noise induces the volunteers to change their voices, and thus Lombard speech in different home scenes can be captured as the training data of the scene acoustic style extractor.

[0078] To improve the training effect, the acoustic features obtained in the noisy home scene can be as rich as possible. For example, by increasing the number of volunteers, the amount of speech of each volunteer in each noisy home scene, etc., the Lombard speech captured in each noisy home scene can be increased, thereby improving the richness of the acoustic features obtained in the noisy home scene.

[0079] By obtaining Lombard speech corresponding to the home scene as richly and comprehensively as possible, extracting acoustic features, obtaining the acoustic features of the home scene in a noisy acoustic environment, and inputting them into the scene acoustic style extractor for training, the accuracy of the training style embedding information extracted by the scene acoustic style extractor can be effectively improved. Correspondingly, the scene style embedding information corresponding to the home scene determined by using the trained scene acoustic style extractor is more accurate and reliable.

[0080] In this embodiment, the acoustic features in a noisy acoustic environment can be various. For example, the acoustic features can be spectral envelope, fundamental frequency (F0) of the sound, acoustic energy, spectrogram such as Linear spectrogram, Mel Frequency Cepstrum Coefficient (MFCC), Mel spectrogram, etc.

[0081] In one implementation, the training style embedding information can be based only on Lombard speech. For example, for each home scene, collect Lombard speech in the noisy home scene and extract acoustic features to obtain the acoustic features of the home scene in a noisy acoustic environment, and input the acoustic features in the noisy acoustic environment into the scene acoustic style extractor to extract the embedded representation (training style embedding information) of the corresponding scene acoustic style features.

[0082] The scene acoustic style extractor can have various implementation manners. Exemplarily, please refer to Figure 5 and Figure 6, the scene acoustic style extractor may include: a reference encoder, a first attention module, and a fully connected layer. Among them, in S220, the step of inputting the acoustic features in a noisy acoustic environment into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene can be implemented through S221, S222, and S223.

[0083] S221, input the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information.

[0084] S222, input the reference embedding information into the first attention module to obtain attention weights.

[0085] S223, input the attention weights into the fully connected layer to obtain the training style embedding information.

[0086] Based on this method, by inputting as rich and comprehensive acoustic features of the home scene in a noisy acoustic environment as possible into the scene acoustic style extractor, the training style embedding information can be obtained, and the training of the scene acoustic style extractor can be realized.

[0087] In another implementation, the training style embedding information can be obtained based on Lombard speech and the scene type information output by the first scene classification model.

[0088] Exemplarily, in the case of inputting the acoustic features of the home scene in a noisy acoustic environment into the scene acoustic style extractor, the environmental acoustic features of the home scene can also be obtained. Input the environmental acoustic features of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene, and input the scene type information into the scene acoustic style extractor for fusion, that is, introduce the scene type information for the learning of the scene acoustic style extractor.

[0089] To improve the convenience and efficiency of fusion, the scene type information can be input into the scene acoustic style extractor in the form of weights for fusion.

[0090] For example, when the first scene classification model is a VGG16 network, the scene type information is the first scene type weight corresponding to the Softmax probability value output by the VGG16 network. The step of inputting the environmental acoustic features of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene may include: inputting the environmental acoustic features of the home scene into the VGG16 network to obtain the Softmax probability value; determining the first scene type weight corresponding to the Softmax probability value; using the first scene type weight as the scene type information.

[0091] Among them, the corresponding relationship between the Softmax probability value and the weight of the first scene type can be preset. Correspondingly, the step of determining the weight of the first scene type corresponding to the Softmax probability value may include: determining the weight of the first scene type corresponding to the Softmax probability value according to the corresponding relationship between the Softmax probability value and the weight of the first scene type.

[0092] Of course, in the case where the scene classification model is a VGG16 network, the label value can also be obtained.

[0093] Please refer to Figure 7 and Figure 8 , in the case where the scene acoustic style extractor includes a reference encoder, a first attention module, and a fully connected layer, Figure 4 The step S220 shown in, that is, inputting the acoustic features in the noisy acoustic environment into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene can be implemented by S224, S225, and S226.

[0094] S224, inputting the acoustic features of the home scene in the noisy acoustic environment into the reference encoder to obtain the reference embedding information.

[0095] S225, inputting the reference embedding information into the first attention module to obtain the attention weight.

[0096] In the case of using the weight of the first scene type as the scene type information, correspondingly, the step of inputting the scene type information into the scene acoustic style extractor for fusion is implemented by S226.

[0097] S226, inputting the attention weight and the weight of the first scene type into the fully connected layer for weighting to obtain the training style embedding information.

[0098] For another example, in the case where the first scene classification model is a ResNet network, the scene type information is the weight of the second scene type corresponding to the label value output by the ResNet network. The step of inputting the environmental acoustic features of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene may include: inputting the environmental acoustic features of the home scene into the ResNet network to obtain the label value of the home scene; determining the weight of the second scene type corresponding to the label value; using the weight of the second scene type as the scene type information.

[0099] Among them, the corresponding relationship between the label value and the weight of the second scene type can be preset. Correspondingly, the step of determining the weight of the second scene type corresponding to the label value may include: determining the weight of the second scene type corresponding to the label value according to the corresponding relationship between the label value and the weight of the second scene type.

[0100] Please refer toFigure 9 and Figure 10 When the scene acoustic style extractor includes a reference encoder, a first attention module, and a fully connected layer, step S220 of inputting the acoustic features in a noisy acoustic environment into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene can be implemented through S227, S228, and S229.

[0101] S227: Input the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information.

[0102] S228: Input the reference embedding information into the first attention module to obtain attention weights.

[0103] When the second scene type weight is used as the scene type information, correspondingly, the step of inputting the scene type information into the scene acoustic style extractor for fusion is implemented through S229.

[0104] S229: Input the attention weights and the second scene type weights into the fully connected layer for weighting to obtain the training style embedding information.

[0105] The above process for obtaining the training style embedding information and the implementation structure of the scene acoustic style extractor are only examples. The process for obtaining the training style embedding information and the implementation structure of the scene acoustic style extractor can also be other. This embodiment does not give examples one by one here. Based on the above process, the training of the scene acoustic style extractor can be achieved.

[0106] In the above example description of the process for obtaining the training style embedding information, the step of inputting the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information may include: inputting the acoustic features of variable-length Lombardspeech (the acoustic features of the home scene in a noisy acoustic environment) into the reference encoder, and the reference encoder compresses the acoustic features of variable-length Lombard speech into fixed-length reference embedding information such as a vector of a fixed size. This reference embedding information is used to encode the overall acoustic style of an audio segment.

[0107] The reference encoder can have various implementation structures. This embodiment gives the following examples.

[0108] For example, as Figure 11 shown, the reference encoder may include a convolutional neural network (Convolutional Neural Networks, CNN), a bidirectional long short-term memory recurrent neural network (including: a backward LSTM and a forward LSTM), and a mapping layer.

[0109] The acoustic features of a home scene in a noisy acoustic environment are input into a CNN to extract further acoustic features. The further acoustic features are input into a bidirectional long short-term memory recurrent neural network to obtain context-related acoustic features. The context-related acoustic features are input into a mapping layer to output fixed-length reference embedding information, encoding the overall acoustic style of an audio segment. Such as the Lombard speech acoustic style of the user in each home scene.

[0110] It can be understood that if different home scenes with similar environmental acoustic features are distinguished, correspondingly, the reference encoder processes the acoustic features of different home scenes with similar environmental acoustic features in a noisy acoustic environment, and can achieve refined processing of the acoustic styles corresponding to different home scenes with similar environmental acoustic features.

[0111] For another example, as Figure 12 shown, the reference encoder may include a first pre-trained model output layer, a CNN, and a mapping layer. The first pre-trained model may refer to the implementation of Audio word2vec, which is mainly implemented by a sequence-to-sequence autoencoder. The main structure in Audio word2vec is a Recurrent Neural Network (RNN). The first pre-trained model output layer outputs pre-trained vectors. The purpose of using Audio word2vec is to obtain a better expression of acoustic features. Then, through two to three convolutional neural networks, further acoustic features are extracted, and then through two to three fully connected networks as the mapping layer, mapped to a predefined dimension.

[0112] For another example, as Figure 13 shown, the reference encoder may include a second pre-trained model output layer, a bidirectional long short-term memory recurrent neural network, and a mapping layer. The second pre-trained model may refer to the implementation of an unsupervised pre-trained model (wav2vec), which is mainly implemented by a CNN-based encoder. The second pre-trained model output layer outputs pre-trained vectors. The purpose of using Wave2Vec is to obtain a better expression of acoustic features. Then, through the bidirectional long short-term memory recurrent neural network, further context-related acoustic features are extracted, and then through two to three fully connected networks as the mapping layer, mapped to a predefined dimension.

[0113] Using the first pre-trained model and the second pre-trained model can obtain more expressions of sound features.

[0114] The first attention module can obtain attention weights in various ways. Exemplarily, an acoustic style marker set can be formed by permuting and combining attributes such as pitch, sound intensity, speech rate, emotion, etc. During the process of the first attention module obtaining attention weights, a group of acoustic style markers is randomly selected from the acoustic style marker set for embedding. Each acoustic style marker can include one or more of the attributes such as pitch, sound intensity, speech rate, emotion, etc. The number of embedded acoustic style markers can be flexibly set. For example, it can be set to one, three, five, ten, etc., to represent a small number of different acoustic dimensions in the training data of the scene acoustic style extractor, such as one or more of the attributes like pitch, sound intensity, speech rate, emotion, etc. Using the reference embedding information as the query information of the first attention module, the first attention module learns the similarity metric between the reference embedding information and each acoustic style marker in the embedded group of acoustic style markers. The first attention module then outputs a set of combined weights (attention weights), and these combined weights represent the contribution of each acoustic style marker to the reference embedding information. Correspondingly, the training style embedding information output by the scene acoustic style extractor may be information related to any one or more of the attributes such as pitch, sound intensity, speech rate, emotion, etc. For example, in the case where a certain home scene is the living room, the training style embedding information corresponding to the living room may include an increased pitch, an increased sound intensity, and an increased speech rate.

[0115] In this embodiment, the training style embedding information can have various presentation forms. For example, the training style embedding information can be specific values assigned to attributes such as pitch, sound intensity, speech rate, emotion, etc. For another example, the training style embedding information can be the adjustment values of attributes such as pitch, sound intensity, speech rate, emotion, etc. based on a certain normal speech. Among them, the normal speech can be the speech uttered in a quiet scene.

[0116] Please refer to Figure 14 , in the case of introducing scene type information such as Softmax probability values or label values to obtain training style embedding information and training the scene acoustic style extractor, the Softmax probability values and label values can be processed into information that can be recognized and used by the fully connected layer through a style embedding regulator. For example, the first scene type weight corresponding to the Softmax probability value is determined through the style embedding regulator, and the second scene type weight corresponding to the label value is determined. In this case, the home scene to which the acoustic features in the noisy acoustic environment belong is known.

[0117] In the case where the scene type information is a tag value, the fully connected layer can multiply the second scene type weight corresponding to the tag value by the attention weight, and only retain one acoustic style marker as the training style embedding information corresponding to the corresponding home scene, so as to determine which attribute has a greater influence on a certain acoustic style marker in this home scene. During training, manual definition of multiple groups of weight parameters can be supported and directly called during inference, thereby improving the training efficiency.

[0118] In the case where the scene type information is a Softmax probability value, the fully connected layer can multiply the first scene type weight corresponding to the softmax probability value by the attention weight to balance the contribution of each embedded acoustic style marker to the reference embedding information again, so that the scene acoustic style extractor can not only learn the acoustic style of Lombard speech in the corresponding home scene, but also learn the influence of different home scenes on the acoustic style.

[0119] To improve the acquisition efficiency of the training style embedding information and accelerate convergence, in another implementation manner, a second scene classification model can also be called. The second scene classification model can indicate the home scene type to which the acoustic features in a noisy acoustic environment belong. By inputting the acoustic features of the home scene in the noisy acoustic environment into the second scene classification model, the scene type indication information corresponding to the home scene can be obtained, so that the training feedback of the scene acoustic style extractor can be performed according to the scene type indication information. For example, during the process of inputting the acoustic features in the noisy acoustic environment into the scene acoustic style extractor to train the scene acoustic style extractor, the acoustic features in the noisy acoustic environment are input into the second scene classification model, and the second scene classification model outputs the scene type indication information, and the scene type indication information is transmitted to the scene acoustic style extractor to perform training feedback on the scene acoustic style extractor, thereby assisting in convergence and improving the acquisition efficiency of the training style embedding information. In this embodiment, when the second scene classification model is called to perform training feedback on the scene acoustic style extractor, the acoustic features in the noisy acoustic environment input into the second scene classification model and the acoustic features in the noisy acoustic environment input into the scene acoustic style extractor can be the same or different, as long as the acoustic features of the home scene input into the scene acoustic style extractor and the acoustic features of the home scene input into the second scene classification model in the noisy acoustic environment correspond to the same home scene.

[0120] The first scene classification model and the second scene classification model can be trained in various ways. Please refer to Figure 15 for the flowchart of a method for training the first scene classification model provided by the embodiment of the present invention, which can be performed by Figure 1The electronic device 100 shown performs, for example, and can be performed by the processor 120 in the electronic device 100. This first scene classification model training method includes S310 and S320.

[0121] S310, obtaining the environmental acoustic features of multiple home scenes.

[0122] S320, inputting the environmental acoustic features of multiple home scenes into a scene classification network for training to obtain scene type information corresponding to each home scene.

[0123] Among them, home scenes can be flexibly divided, and the environmental acoustic features of home scenes can be collected. For example, the kitchen, living room, and bathroom where noisy sounds often occur can be selected as home scenes respectively and the environmental acoustic features of the home scenes can be collected. For another example, a home exercise gym, a multimedia audio-visual room, a study, etc. can also be selected as home scenes respectively and the environmental acoustic features of the home scenes can be collected. Through the fine division of home scenes and using the environmental acoustic features of different home scenes to train the scene classification network, scene type information corresponding to each home scene is obtained, and based on the scene type information, the judgment of the home scene type can be realized.

[0124] To improve the training effect, the environmental acoustic features of each home scene obtained can be as rich as possible. For example, the environmental acoustic features of each home scene can include the individual environmental acoustic features of each home device located in this home scene, and the environmental acoustic features of at least pairwise permutations and combinations of each home device located in this home scene. Exemplarily, when the home scene is the kitchen, the environmental acoustic features of the kitchen can include the individual environmental acoustic features of each home device in the kitchen, such as the range hood, dishwasher, gas stove, pots and pans, etc., and the environmental acoustic features of permutations and combinations of two, three, four, etc. of the range hood, dishwasher, gas stove, pots and pans, etc. in the kitchen.

[0125] By obtaining the environmental acoustic features of home scenes as richly and comprehensively as possible and inputting them into the scene classification network for training, the sensitivity and reliability of the first scene classification model obtained through training can be effectively improved. For example, by inputting the individual environmental acoustic features of home devices into the scene classification network for training, the type of home scene can be directly determined based on the environmental acoustic features of certain specific home devices subsequently. Exemplarily, based on the acoustic feature of the toilet flushing sound, the home scene can be directly determined as the bathroom. For another example, by inputting the environmental acoustic features of various permutations and combinations of each home device in each home scene into the scene classification network for training, reliable distinction between different home scenes with similar environmental acoustic features can be achieved. Exemplarily, both the kitchen and the bathroom may have the acoustic feature of running water mixed in. However, combining the acoustic feature of the sound of stir-frying in a wok can determine that the home scene is the kitchen.

[0126] Please refer to Figure 16 In one implementation, in S310, environmental acoustic features of multiple home scenarios are obtained, which can be achieved through S311 and S312.

[0127] S311, obtain environmental acoustic data of multiple home scenarios.

[0128] S312, perform acoustic feature extraction on the environmental acoustic data to obtain environmental acoustic features of multiple home scenarios.

[0129] Among them, acoustic feature extraction can be performed in multiple ways. For example, an acoustic feature extraction module can be used to perform acoustic feature extraction on the environmental acoustic data. The environmental acoustic data in this embodiment is the acoustic data generated during the use of each home device in the home scenario and does not include human voices.

[0130] The scene classification network can be flexibly selected. For example, it can be a set deep learning network. Exemplarily, the scene classification network can be a deep convolutional neural network, such as the VGG16 network, the ResNet network, etc. According to the different scene classification networks, different scene type information can be obtained. Exemplarily, when the scene classification network is the VGG16 network, in S320, the environmental acoustic features are input into the VGG16 network, and the obtained scene type information is the Softmax probability value (of course, a label value such as one-hot encoding can also be obtained). When the scene classification network is the ResNet network, in S320, the environmental acoustic features are input into the ResNet network, and the obtained scene type information is a label value such as one-hot encoding (of course, the obtained scene type information can also be the softmax probability value).

[0131] Based on the above design, according to the label value representing the home scene type, the specific home scene type to which the environmental acoustic features of the home scene belong can be determined, or according to the softmax probability value representing the home scene belonging to each type, different weight combinations of the home scene style can be determined, so as to realize the refinement of the home scene recognition granularity, improve the accuracy of home scene recognition, and can also better "guide" the extraction of scene style embedding.

[0132] Based on the refined classification of the home scene by the first scene classification model, the scene acoustic style under the home scene can be divided more meticulously. And it can also better "guide" the extraction of scene style embedding. Exemplarily, when the home scene includes three types: kitchen, living room, and bathroom, the scene acoustic style extractor can perform scene acoustic style extraction for the kitchen, living room, and bathroom respectively in combination with the first scene classification model.

[0133] To more clearly illustrate the first scenario classification model training method in the embodiments of the present invention, the following scenario is taken as an example for illustration.

[0134] In the case where the home scene includes three types: kitchen, living room, and bathroom, collect the environmental acoustic data generated separately by each home appliance in the kitchen, as well as the environmental acoustic data generated by two permutations, three permutations, four permutations, etc. Then, take all the collected environmental acoustic data in the kitchen as the first acoustic data set. Similarly, collect the environmental acoustic data generated separately by each home appliance in the living room, as well as the environmental acoustic data generated by two permutations, three permutations, four permutations, etc., and take all the collected environmental acoustic data in the living room as the second acoustic data set. Also, collect the environmental acoustic data generated separately by each home appliance in the bathroom, as well as the environmental acoustic data generated by two permutations, three permutations, four permutations, etc., and take all the collected environmental acoustic data in the bathroom as the third acoustic data set.

[0135] Extract the acoustic features of the environmental acoustic data in the first acoustic data set to obtain the first acoustic feature set. Extract the acoustic features of the environmental acoustic data in the second acoustic data set to obtain the second acoustic feature set. Extract the acoustic features of the environmental acoustic data in the third acoustic data set to obtain the third acoustic feature set.

[0136] Set the scenario type information corresponding to the kitchen as the first type, the scenario type information corresponding to the living room as the second type, and the scenario type information corresponding to the bathroom as the third type.

[0137] When inputting the acoustic features in the first acoustic feature set into the scenario classification network, set the first type as the target output of the scenario classification network. When inputting the acoustic features in the second acoustic feature set into the scenario classification network, set the second type as the target output of the scenario classification network. When inputting the acoustic features in the third acoustic feature set into the scenario classification network, set the third type as the target output of the scenario classification network. Based on this, train the scenario classification network until the convergence condition is met, and then the required first scenario classification model can be obtained.

[0138] The convergence condition can be set flexibly. For example, it can be that the accuracy rate of obtaining the scenario type information corresponding to each home scene reaches a preset value. Exemplarily, the acoustic features in the first acoustic feature set, the second acoustic feature set, and the third acoustic feature set can be divided into training data and test data. Then, train the scenario classification network based on the training data until the accuracy rate of obtaining the target output after inputting the test data reaches the preset value, and it is determined that the convergence condition is met, and the required first scenario classification model is obtained. Another example is that the number of training times reaches a set amount. This embodiment does not limit this.

[0139] After obtaining the first scene classification model by using the above scene classification model training method, based on the first scene classification model, the weight combination of the specific home scene or home scene style to which the environmental acoustic features belong can be identified. For example, if the environmental acoustic features of a certain home scene are input into the scene classification model, and the scene classification model outputs the first type, it can be determined that the environmental acoustic features belong to the kitchen according to the first type.

[0140] After inputting the environmental acoustic features of multiple home scenes into the trained first scene classification model, the first scene classification model outputs scene type related information, such as label values or softmax probability values. The scene type information output by the first scene classification model can be used as the input of the scene acoustic style extractor for combined training and use of the scene acoustic style extractor.

[0141] The embodiment of the present invention also provides a second scene classification model training method, including: obtaining the acoustic features of multiple home scenes in a noisy acoustic environment, inputting the acoustic features of multiple home scenes in a noisy acoustic environment into a scene classification network for training, and obtaining scene type indication information corresponding to each home scene.

[0142] The second scene classification model uses the acoustic features of multiple home scenes in a noisy acoustic environment as input. Correspondingly, the second scene classification model can indicate the home scene type to which the acoustic features in the noisy acoustic environment belong.

[0143] Since the difference between the training processes of the second scene classification model and the first scene classification model lies only in the different input data, due to the different input data, the output data is different. The scene type indication information output by the second scene classification model is used as the training feedback of the scene acoustic style extractor. For example, the fully connected layer of the scene acoustic style extractor is trained and fed back according to the scene type indication information to accelerate convergence. The scene type information output by the first scene classification model is used for weighting. For example, the scene type information and the attention weight are input into the fully connected layer of the scene acoustic style extractor for weighting. The optional structure and training principle of the second scene classification model can refer to the relevant description of the above first scene classification model, so it will not be elaborated here.

[0144] In the process of inputting the acoustic features of a home scene in a noisy acoustic environment into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene, the first scene classification model can be selected to be called, or the second scene classification model can be selected to be called, or both the first scene classification model and the second scene classification model can be selected to be called, or neither the first scene classification model nor the second scene classification model can be called.

[0145] Please refer to Figure 17, the predicted features of the voice to be synthesized can be obtained through S410 and S420.

[0146] S410, obtain the information of the text to be output and the training style embedding information corresponding to the home scene.

[0147] S420, input the information of the text to be output and the training style embedding information into the acoustic feature prediction network to obtain the predicted features of the voice to be synthesized corresponding to the home scene.

[0148] The training style embedding information corresponding to the home scene can be obtained through the aforementioned scene acoustic style extractor, which will not be elaborated here.

[0149] In this embodiment, the training style embedding information corresponding to the home scene is ingeniously introduced as a new dimension for voice synthesis consideration. The information of the text to be output and the training style embedding information are both input into the acoustic feature prediction network to obtain the predicted features of the voice to be synthesized corresponding to the home scene, thereby realizing the training of the acoustic feature prediction model. The predicted features of the voice to be synthesized have the Lombard speech acoustic style. Based on the two dimensions of the information of the text to be output and the training style embedding information, the reliable prediction of the acoustic features corresponding to the home scene can be realized, and the predicted acoustic features with the Lombard speech acoustic style can be output.

[0150] In one implementation, the voice uttered by the user in a quiet scene can be collected as normal voice, and the acoustic feature prediction network can be jointly trained with the normal voice and the training style embedding information to obtain the predicted features of the voice to be synthesized corresponding to the home scene. For example, when the training style embedding information is to increase the pitch, increase the sound intensity, and increase the speech rate, the acoustic feature prediction network is jointly trained with the normal voice and the training style embedding information, and the predicted features of the voice to be synthesized obtained may include increasing the volume, sound intensity, and speech rate by set values respectively on the basis of the normal voice.

[0151] The acoustic feature prediction network can have multiple implementation manners. In order to improve the naturalness of the predicted features of the voice to be synthesized, Tacotron can be selected as the acoustic feature prediction network.

[0152] Please refer to Figure 18, in one implementation, the acoustic feature prediction network may include: an encoder, a second attention module, and a decoder. Correspondingly, for S420, the step of inputting the text information to be output and the training style embedding information into the acoustic feature prediction network to obtain the predicted features of the sound to be synthesized corresponding to the home scene can be implemented in the following way: input the text information to be output into the encoder to obtain character embedding information of a fixed length; input the character embedding information of the fixed length and the training style embedding information into the second attention module to obtain the aligned acoustic features and character information; input the aligned acoustic features and character information into the decoder to obtain the predicted features of the sound to be synthesized.

[0153] When Tacotron is selected as the acoustic feature prediction network, inputting the aligned acoustic features and character information into the decoder, the predicted features of the sound to be synthesized output by the decoder can be a linear spectrum or a mel spectrum or other acoustic features applicable to a vocoder.

[0154] After obtaining the predicted features of the sound to be synthesized, a speech synthesis device can be used to synthesize the predicted features of the sound to be synthesized to obtain the speech data to be output. The speech synthesis device can be a vocoder. Correspondingly, a vocoder can be used as the speech synthesis device to synthesize the predicted features of the sound to be synthesized output by the acoustic feature prediction network, and input the predicted features of the sound to be synthesized into the vocoder to obtain the speech data to be output.

[0155] Through the training of the speech synthesis model, after inputting the acoustic features of the home scene in a noisy acoustic environment and the response text content information into the trained speech synthesis model, the speech synthesis model can output the output response speech data containing the response text content information, and the output response speech data is Lombard speech with good naturalness and clarity.

[0156] Please refer to Figure 19 , which is a schematic flowchart of a speech synthesis method provided by an embodiment of the present invention, and can be executed by Figure 1 the electronic device 100 shown, for example, can be executed by the processor 120 in the electronic device 100. The speech synthesis method includes S510, S520, and S530.

[0157] S510, determine the scene style embedding information corresponding to the home scene; wherein, the scene style embedding information represents the scene acoustic style corresponding to the home scene.

[0158] S520, determine the predicted acoustic features corresponding to the home scene according to the response text content information and the scene style embedding information.

[0159] S530 synthesizes the predicted acoustic features to obtain output response speech data.

[0160] In one implementation, the electronic device can be a smart device in a smart home scenario. The output response speech data synthesized by the smart device through S510 to S530 is Lombard speech with the Lombard effect. The smart device outputs the output response speech data in a noisy home scenario so that it can be clearly transmitted to the user, ensuring smooth interaction. In another implementation, the electronic device can be a server, which can be communicatively connected to the smart device in the smart home scenario. The output response speech data synthesized by the server through S510 to S530 is Lombard speech with the Lombard effect. The server sends the synthesized output response speech data to the smart device, and the smart device plays the output response speech data in a noisy home scenario so that it can be clearly transmitted to the user, ensuring smooth interaction.

[0161] There are various implementation manners for S510 to S530. For example, the scene style embedding information corresponding to the home scenario can be determined by a trained scene acoustic style extractor. Another example is that the predicted acoustic features corresponding to the home scenario can be determined by a trained acoustic feature prediction model.

[0162] Please refer to Figure 20 In one implementation manner, the electronic device can call a trained scene acoustic style extractor to obtain the acoustic features of the home scenario in a noisy acoustic environment, input the acoustic features of the home scenario in the noisy acoustic environment into the scene acoustic style extractor to obtain the scene style embedding information corresponding to the home scenario. The electronic device can call a trained acoustic feature prediction model, use the response text content information and the scene style embedding information as the input of the acoustic feature prediction model to obtain the predicted acoustic features corresponding to the home scenario, and then synthesize the output response speech data.

[0163] Please refer to Figure 21, In another implementation, the electronic device can call the trained scene acoustic style extractor and the first scene classification model, input the environmental acoustic features of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene, and input the scene type information into the scene acoustic style extractor for fusion. For example, the scene type information and the acoustic features of the home scene in a noisy acoustic environment are jointly used as the input of the scene acoustic style extractor, and the scene acoustic style extractor outputs the scene style embedding information corresponding to the home scene. The electronic device can call the trained acoustic feature prediction model, use the response text content information and the scene style embedding information as the input of the acoustic feature prediction model to obtain the predicted acoustic features corresponding to the home scene, and then synthesize and output the response voice data.

[0164] In this embodiment, the training processes of the scene acoustic style extractor, the acoustic feature prediction model, the first scene classification model, and the second scene classification model can refer to the corresponding descriptions in the speech synthesis model training method, which will not be elaborated here. The trained scene acoustic style extractor, acoustic feature prediction model, first scene classification model, and second scene classification model can run independently or be combined and run in different combinations according to requirements.

[0165] For example, during the speech synthesis process, the trained scene acoustic style extractor and acoustic feature prediction model can be called. Please refer to Figure 22 , The scene acoustic style extractor can include: a reference encoder, a first attention module, and a fully connected layer. In the case where only the acoustic features of the home scene in a noisy acoustic environment are used as the input of the scene acoustic style extractor, the scene style embedding information corresponding to the home scene can be obtained in the following way: input the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information; input the reference embedding information into the first attention module to obtain attention weights; input the attention weights into the fully connected layer to obtain the scene style embedding information. Use the response text content information and the scene style embedding information as the input of the acoustic feature prediction model, so as to obtain the predicted acoustic features corresponding to the home scene, synthesize the predicted acoustic features, and then obtain the output response voice data.

[0166] Another example is that during the speech synthesis process, the trained scene acoustic style extractor, the first scene classification model, and the acoustic feature prediction model can be called. Please refer to Figure 23, the scene acoustic style extractor may include: a reference encoder, a first attention module, and a fully connected layer. When the scene type information and the acoustic features of the home scene in a noisy acoustic environment are jointly used as the input of the scene acoustic style extractor, correspondingly, the scene style embedding information corresponding to the home scene can be obtained in the following way: input the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information; input the reference embedding information into the first attention module to obtain attention weights; input the attention weights and the scene type information into the fully connected layer for weighting to obtain the scene style embedding information. Use the response text content information and the scene style embedding information as the input of the acoustic feature prediction model to obtain the predicted acoustic features corresponding to the home scene, and synthesize the predicted acoustic features to further obtain the output response speech data.

[0167] When the first scene classification model is a VGG16 network, the environmental acoustic features of the home scene can be input into the VGG16 network to obtain Softmax probability values; determine the first scene type weights corresponding to the Softmax probability values; use the first scene type weights as the scene type information. Among them, the first scene type weights corresponding to the Softmax probability values can be determined according to the corresponding relationship between the Softmax probability values and the first scene type weights.

[0168] Correspondingly, input the attention weights and the first scene type weights into the fully connected layer for weighting to obtain the scene style embedding information.

[0169] When the first scene classification model is a ResNet network, the environmental acoustic features of the home scene can be input into the ResNet network to obtain the label values of the home scene; determine the second scene type weights corresponding to the label values; use the second scene type weights as the scene type information. Among them, the second scene type weights corresponding to the label values can be determined according to the corresponding relationship between the label values and the second scene type weights.

[0170] Correspondingly, input the attention weights and the second scene type weights into the fully connected layer for weighting to obtain the scene style embedding information.

[0171] In this embodiment, the acoustic features of the home scene in a noisy acoustic environment are used as the input of the scene acoustic style extractor, or the acoustic features of the home scene in a noisy acoustic environment and the scene type information output by the above first scene classification model are jointly used as the input of the scene acoustic style extractor, and the scene acoustic style extractor thus outputs the scene style embedding information. The scene style embedding information output by the scene acoustic style extractor can be used as the input of the following acoustic feature prediction model.

[0172] Please refer to Figure 24, the acoustic feature prediction model may include an encoder, a second attention module, and a decoder. Correspondingly, taking the response text content information and the scene style embedding information as the input of the acoustic feature prediction model, obtaining the predicted acoustic features corresponding to the home scene can be achieved in the following manner: inputting the response text content information into the encoder to obtain fixed-length character embedding information; inputting the fixed-length character embedding information and the scene style embedding information into the second attention module to obtain the aligned acoustic features and character information; inputting the aligned acoustic features and character information into the decoder to obtain the predicted acoustic features corresponding to the home scene.

[0173] In this embodiment, taking the response text content information and the scene style embedding information as the input of the acoustic feature prediction model, the acoustic feature prediction model thus outputs the predicted acoustic features. Based on the predicted acoustic features, the output response speech data with the Lombard effect can be synthesized, so that the output response speech data played by the electronic device in a noisy home scene can be reliably received by the user, thereby ensuring the smoothness of human-computer interaction in the smart home scene and better meeting the actual application requirements.

[0174] In one implementation, a speech synthesis device can be used to synthesize the output response speech data. As Figure 25 shown, a vocoder can be used as the speech synthesis device to synthesize the predicted acoustic features output by the acoustic feature prediction model. Inputting the predicted acoustic features into the vocoder to obtain the output response speech data containing the response text content information. In this embodiment, the vocoder can be flexibly selected. For example, when the predicted acoustic feature is Linear Spectrum, Griffin-Lim can be selected as the vocoder for converting the spectrum to waveform, or WaveNet can be selected as the vocoder. Another example is that when the predicted acoustic feature is Mel Spectrum, WaveNet can be selected as the vocoder.

[0175] The above describes the optional embodiments of the speech synthesis method in the embodiments of the present invention. In other implementations, based on the same design concept: by "imitating" the communication mode in which humans actively change their vocalization methods under the Lombard effect, for different home scenes, synthesize speech data with corresponding scene acoustic features for broadcasting, and ensure the smoothness of voice interaction with users in a noisy home environment by synthesizing Lombard speech with good naturalness and clarity. There may be other implementations of the speech synthesis method in the embodiments of the present invention.

[0176] In another implementation, please refer to Figure 26, the Lombard speech can be input into the Lombard speech generation model, and the Lombard speech generation model learns based on the Lombard speech, so that the Lombard speech generation model can directly obtain the output response speech data according to the response text content information. Based on this kind of Lombard speech generation model, the intelligent device does not need to train and call the first scene classification model, the second scene classification model, the scene acoustic style extractor and the acoustic feature prediction model, and directly inputs the response text content information into the Lombard speech generation model, then the output response speech data corresponding to the home scene can be obtained.

[0177] In this embodiment, a speech synthesis device can be used to synthesize the predicted acoustic features to obtain the output response speech data. In one implementation, the speech synthesis device can be a vocoder, and the predicted acoustic features are input into the vocoder, and then the output response speech data is obtained.

[0178] To more clearly illustrate the speech synthesis method in the embodiments of the present invention, the following scenario is taken as an example for illustration.

[0179] The electronic device is an intelligent device. The intelligent device is located in a home scene and can perform voice interaction with the user, such as playing the output response speech data. The intelligent device is loaded with a first scene classification model, a second scene classification model, a scene acoustic style extractor and an acoustic feature prediction model.

[0180] The intelligent device determines whether the user has made a pre-configuration. If it is determined that the user has made a pre-configuration, the output response speech data is generated according to the pre-configuration; if it is determined that the user has not made a pre-configuration, the intelligent device receives the surrounding sounds and determines whether the received sound is only the environmental sound of the home scene or the user's voice in a noisy acoustic environment.

[0181] If it is determined that the received sound is only the environmental sound of the home scene, the first scene classification model is called, and the environmental sound of the home scene is input into the first scene classification model to obtain the scene type information corresponding to the home scene, such as the label value or the softmax probability value, for classifying the home scene.

[0182] If it is determined that the received sound is the user's voice in a noisy acoustic environment, the intelligent device calls the scene acoustic style extractor and the second scene classification model. When the first scene classification model has not obtained the scene type information corresponding to the home scene, the user's voice in the noisy acoustic environment is input into the scene acoustic style extractor and the second scene classification model to generate the corresponding scene style embedding information. When the first scene classification model has obtained the scene type information corresponding to the home scene, the user's voice in the noisy acoustic environment and the scene type information corresponding to the home scene are input into the scene acoustic style extractor to generate the corresponding scene style embedding information.

[0183] Before inputting the user speech in a noisy acoustic environment into the scene acoustic style extractor, noise reduction processing can also be performed, and the denoised sound is input into the scene acoustic style extractor to further improve the accuracy of extracting the scene style embedding information.

[0184] When the intelligent device detects that the user issues a question, it determines whether the scene style embedding information corresponding to the home scene has been obtained. If it is determined that the scene style embedding information corresponding to the home scene has been obtained, it calls the acoustic feature prediction model trained based on the to-be-output text information and the scene style embedding information, inputs the response text content information and the scene style embedding information corresponding to the home scene into the acoustic feature prediction model, generates the predicted acoustic features of Lombard speech in the corresponding home scene, and synthesizes and outputs the response voice data through a vocoder. If it is determined that the scene style embedding information corresponding to the home scene has not been obtained, it calls the Lombard speech generation model, inputs the response text content information into the Lombard speech generation model, generates the predicted acoustic features of Lombard speech in the corresponding home scene, and synthesizes and outputs the response voice data through a vocoder. The intelligent device broadcasts the synthesized output response voice data, so as to be able to interact with the user in a voice with the acoustic style in the corresponding home scene. The broadcast voice is Lombard speech with better naturalness and clarity, so as to ensure the smoothness of the voice interaction. It can be understood that the above voice interaction process can be repeatedly executed according to user needs.

[0185] To execute the corresponding steps in the above embodiments and each possible manner, an implementation manner of a voice synthesis device is given below. Please refer to Figure 27 , Figure 27 FIG. 140 is a functional module diagram of a first voice synthesis device 140 provided by an embodiment of the present invention. The first voice synthesis device 140 can be applied to Figure 1 the electronic device 100 shown in FIG. It should be noted that for the first voice synthesis device 140 provided in this embodiment, its basic principle and the generated technical effects are the same as those in the above voice synthesis method embodiment. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above voice synthesis method embodiment. The first voice synthesis device 140 includes an information determination module 141 and a response voice data synthesis module 142.

[0186] The information determination module 141 is configured to determine the scene style embedding information corresponding to the home scene, and determine the predicted acoustic features corresponding to the home scene according to the response text content information and the scene style embedding information, where the scene style embedding information represents the scene acoustic style corresponding to the home scene.

[0187] The response speech data synthesis module 142 is used to synthesize the predicted acoustic features to obtain output response speech data.

[0188] Please refer to Figure 28 , the embodiments of the present invention also provide an implementation manner of the second speech synthesis device 150. The second speech synthesis device 150 includes: a predicted acoustic feature obtaining module 151 and an information synthesis module 152.

[0189] Among them, the predicted acoustic feature obtaining module 151 is used to input the response text content information into an acoustic feature prediction model to obtain predicted acoustic features corresponding to the home scene.

[0190] The information synthesis module 152 is used to synthesize the predicted acoustic features to obtain output response speech data.

[0191] To execute the corresponding steps in the above embodiments and each possible manner, an implementation manner of a scene classification model training device is given below. Please refer to Figure 29 , Figure 29 is a functional module diagram of a scene classification model training device 160 provided by an embodiment of the present invention. The scene classification model training device 160 can be applied to Figure 1 the electronic device 100 shown. It should be noted that for the scene classification model training device 160 provided in this embodiment, its basic principle and the generated technical effects are the same as those in the above embodiment of the scene classification model training method. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above embodiment of the scene classification model training method. The scene classification model training device 160 includes an environmental acoustic feature obtaining module 161 and a scene classification network training module 162.

[0192] Among them, the environmental acoustic feature obtaining module 161 is used to obtain the environmental acoustic features of multiple home scenes.

[0193] The scene classification network training module 162 is used to input the environmental acoustic features of the multiple home scenes into a scene classification network for training to obtain scene type information corresponding to each of the home scenes.

[0194] To execute the corresponding steps in the above embodiments and each possible manner, an implementation manner of a scene acoustic style extractor training device is given below. Please refer to Figure 30 , Figure 30 is a functional module diagram of a scene acoustic style extractor training device 170 provided by an embodiment of the present invention. The scene acoustic style extractor training device 170 can be applied to Figure 1The electronic device 100 shown. It should be noted that for the scene acoustic style extractor training device 170 provided in this embodiment, its basic principle and the technical effects produced are the same as those in the above-mentioned scene acoustic style extractor training method embodiment. For a brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above-mentioned scene acoustic style extractor training method embodiment. The scene acoustic style extractor training device 170 includes a noisy environment acoustic feature acquisition module 171 and a scene acoustic style extractor training module 172.

[0195] Among them, the noisy environment acoustic feature acquisition module 171 is used to acquire the acoustic features of the home scene in a noisy acoustic environment.

[0196] The scene acoustic style extractor training module 172 is used to input the acoustic features of the home scene in the noisy acoustic environment into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene; wherein, the training style embedding information represents the scene acoustic style corresponding to the home scene.

[0197] To execute the corresponding steps in the above-mentioned embodiments and various possible ways, an implementation manner of an acoustic feature prediction model training device is given below. Please refer to Figure 31 , Figure 31 is a functional module diagram of an acoustic feature prediction model training device 180 provided by an embodiment of the present invention. The acoustic feature prediction model training device 180 can be applied to Figure 1 the electronic device 100 shown. It should be noted that for the acoustic feature prediction model training device 180 provided in this embodiment, its basic principle and the technical effects produced are the same as those in the above-mentioned acoustic feature prediction model training method embodiment. For a brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above-mentioned acoustic feature prediction model training method embodiment. The acoustic feature prediction model training device 180 includes a data acquisition module 181 and an acoustic feature prediction model training module 182.

[0198] Among them, the data acquisition module 181 is used to acquire the text information to be output and the scene style embedding information corresponding to the home scene.

[0199] The acoustic feature prediction model training module 182 is used to input the text information to be output and the training style embedding information into the acoustic feature prediction network to obtain the predicted sound features to be synthesized corresponding to the home scene.

[0200] To execute the corresponding steps in the above-mentioned embodiments and various possible ways, an implementation manner of a speech synthesis model training device is given below. Please refer to Figure 32 , Figure 32FIG. 0 is a functional block diagram of a voice synthesis model training apparatus 190 provided by an embodiment of the present invention. The voice synthesis model training apparatus 190 can be applied to Figure 1 the electronic device 100 shown in FIG. It should be noted that the basic principle and the technical effects produced by the voice synthesis model training apparatus 190 provided in this embodiment are the same as those in the above-mentioned voice synthesis model training method embodiment. For a brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above-mentioned voice synthesis model training method embodiment. The voice synthesis model training apparatus 190 includes an information acquisition module 191 and a to-be-output voice data synthesis module 192.

[0201] Among them, the information acquisition module 191 is configured to acquire to-be-output text information and scene style embedding information corresponding to a home scene, and obtain to-be-synthesized sound prediction features corresponding to the home scene according to the to-be-output text information and the scene style embedding information; wherein, the training style embedding information represents the scene acoustic style corresponding to the home scene.

[0202] The to-be-output voice data synthesis module 192 is configured to synthesize the to-be-synthesized sound prediction features to obtain to-be-output voice data.

[0203] The information acquisition module 191 is configured to obtain the training style embedding information corresponding to the home scene through the following steps: acquire the acoustic features of the home scene in a noisy acoustic environment; input the acoustic features of the home scene in the noisy acoustic environment into a scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene.

[0204] The information acquisition module 191 is further configured to: input the environmental acoustic features of the home scene into a first scene classification model to obtain scene type information corresponding to the home scene; input the scene type information into the scene acoustic style extractor for fusion.

[0205] When the first scene classification model is a VGG16 network, the information acquisition module 191 is configured to input the environmental acoustic features of the home scene into the first scene classification model through the following steps to obtain scene type information corresponding to the home scene: input the environmental acoustic features of the home scene into the VGG16 network to obtain a Softmax probability value; determine a first scene type weight corresponding to the Softmax probability value; use the first scene type weight as the scene type information.

[0206] The scene acoustic style extractor includes: a reference encoder, a first attention module, and a fully connected layer; the information acquisition module 191 is configured to input the acoustic features of the home scene in the noisy acoustic environment into the scene acoustic style extractor through the following steps to obtain the training style embedding information corresponding to the home scene: input the acoustic features of the home scene in the noisy acoustic environment into the reference encoder to obtain reference embedding information; input the reference embedding information into the first attention module to obtain attention weights; input the attention weights and the first scene type weight into the fully connected layer for weighting to obtain the training style embedding information.

[0207] The information acquisition module 191 is configured to determine the first scene type weight corresponding to the Softmax probability value through the following steps: determine the first scene type weight corresponding to the Softmax probability value according to the correspondence between the Softmax probability value and the first scene type weight.

[0208] When the first scene classification model is a ResNet network, the information acquisition module 191 is configured to input the environmental acoustic features of the home scene into the first scene classification model through the following steps to obtain the scene type information corresponding to the home scene: input the environmental acoustic features of the home scene into the ResNet network to obtain the label value of the home scene; determine the second scene type weight corresponding to the label value; use the second scene type weight as the scene type information.

[0209] The scene acoustic style extractor includes: a reference encoder, a first attention module, and a fully connected layer; the information acquisition module 191 is configured to input the acoustic features of the home scene in the noisy acoustic environment into the scene acoustic style extractor through the following steps to obtain the training style embedding information corresponding to the home scene: input the acoustic features of the home scene in the noisy acoustic environment into the reference encoder to obtain reference embedding information; input the reference embedding information into the first attention module to obtain attention weights; input the attention weights and the second scene type weight into the fully connected layer for weighting to obtain the training style embedding information.

[0210] The information acquisition module 191 is configured to determine the second scene type weight corresponding to the label value through the following steps: determine the second scene type weight corresponding to the label value according to the correspondence between the label value and the second scene type weight.

[0211] The information acquisition module 191 is further configured to train the first scene classification model through the following steps: obtain the environmental acoustic features of multiple home scenes; input the environmental acoustic features of the multiple home scenes into a scene classification network for training to obtain the scene type information corresponding to each home scene.

[0212] The scene acoustic style extractor includes: a reference encoder, a first attention module, and a fully connected layer; the information acquisition module 191 is configured to input the acoustic features of the home scene in the noisy acoustic environment into the scene acoustic style extractor through the following steps to obtain the training style embedding information corresponding to the home scene: input the acoustic features of the home scene in the noisy acoustic environment into the reference encoder to obtain reference embedding information; input the reference embedding information into the first attention module to obtain attention weights; input the attention weights into the fully connected layer to obtain the training style embedding information.

[0213] The information acquisition module 191 is further configured to: input the acoustic features of the home scene in the noisy acoustic environment into a second scene classification model to obtain the scene type indication information corresponding to the home scene; perform training feedback on the scene acoustic style extractor according to the scene type indication information.

[0214] The information acquisition module 191 is further configured to train the second scene classification model through the following steps: obtain the acoustic features of multiple home scenes in the noisy acoustic environment; input the acoustic features of the multiple home scenes in the noisy acoustic environment into a scene classification network for training to obtain the scene type indication information corresponding to each home scene.

[0215] The information acquisition module 191 is configured to obtain the predicted feature of the sound to be synthesized corresponding to the home scene according to the text information to be output and the training style embedding information through the following steps: input the text information to be output and the training style embedding information as the input of an acoustic feature prediction network to obtain the predicted feature of the sound to be synthesized corresponding to the home scene.

[0216] The acoustic feature prediction network includes an encoder, a second attention module, and a decoder; the information acquisition module 191 is configured to input the text information to be output and the training style embedding information as the input of the acoustic feature prediction network through the following steps to obtain the predicted feature of the sound to be synthesized corresponding to the home scene: input the text information to be output into the encoder to obtain character embedding information with a fixed length; input the character embedding information with the fixed length and the training style embedding information into the second attention module to obtain the aligned acoustic feature and character information; input the aligned acoustic feature and character information into the decoder to obtain the predicted feature of the sound to be synthesized corresponding to the home scene.

[0217] The to-be-output speech data synthesis module 192 is used to synthesize the to-be-synthesized voice prediction features through the following steps to obtain the to-be-output speech data: Use a vocoder to synthesize the to-be-synthesized voice prediction features to obtain the to-be-output speech data.

[0218] The to-be-synthesized voice prediction features are Linear Spectrum, and the vocoder is Griffin-Lim or WaveNet; or, the to-be-synthesized voice prediction features are Mel Spectrum, and the vocoder is WaveNet.

[0219] On the above basis, an embodiment of the present invention further provides a computer-readable storage medium, which includes a computer program. When the computer program runs, it controls the electronic device where the computer-readable storage medium is located to execute the above-mentioned speech synthesis model training method.

[0220] In the embodiment of the present invention, the synthesized output response speech data is Lombard speech with good naturalness and clarity, which can ensure the smoothness of voice interaction and improve the user experience.

[0221] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0222] In addition, the functional modules in each embodiment of the present invention can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.

[0223] When the above-mentioned function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0224] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for training a speech synthesis model, characterized in that, it includes: obtaining the text information to be output and the training style embedding information corresponding to the home scene; wherein, the training style embedding information represents the scene acoustic style corresponding to the home scene; obtaining the predicted sound features to be synthesized corresponding to the home scene according to the text information to be output and the training style embedding information; synthesizing the predicted sound features to be synthesized to obtain the speech data to be output; the training style embedding information corresponding to the home scene is obtained through the following steps: obtaining the acoustic features of the home scene in a noisy acoustic environment; inputting the acoustic features of the home scene in the noisy acoustic environment into a scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene; further includes: obtaining the environmental acoustic features of the home scene; inputting the environmental acoustic features of the home scene into a first scene classification model to obtain the scene type information corresponding to the home scene; inputting the scene type information into the scene acoustic style extractor for fusion.

2. The method for training a speech synthesis model according to claim 1, characterized in that, when the first scene classification model is a VGG16 network, the scene type information is the first scene type weight corresponding to the Softmax probability value output by the VGG16 network.

3. The method for training a speech synthesis model according to claim 1, characterized in that, when the first scene classification model is a ResNet network, the scene type information is the second scene type weight corresponding to the label value output by the ResNet network.

4. The method for training a speech synthesis model according to claim 1, characterized in that, further includes: inputting the acoustic features of the home scene in a noisy acoustic environment into a second scene classification model to obtain the scene type indication information corresponding to the home scene; performing training feedback on the scene acoustic style extractor according to the scene type indication information.

5. The method for training a speech synthesis model according to claim 1, characterized in that, the step of obtaining the predicted sound features to be synthesized corresponding to the home scene according to the text information to be output and the training style embedding information includes: using the text information to be output and the training style embedding information as the input of an acoustic feature prediction network to obtain the predicted sound features to be synthesized corresponding to the home scene.

6. A device for training a speech synthesis model, characterized in that, it includes: an information acquisition module, configured to obtain the text information to be output and the training style embedding information corresponding to the home scene, and obtain the predicted sound features to be synthesized corresponding to the home scene according to the text information to be output and the training style embedding information; wherein, the training style embedding information represents the scene acoustic style corresponding to the home scene; the information acquisition module is further configured to obtain the acoustic features of the home scene in a noisy acoustic environment; input the acoustic features of the home scene in the noisy acoustic environment into a scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene; The information acquisition module is further configured to acquire the environmental acoustic features of the home scene; input the environmental acoustic features of the home scene into a first scene classification model to obtain the scene type information corresponding to the home scene; and input the scene type information into the scene acoustic style extractor for fusion. The speech data to be output synthesis module is configured to synthesize the predicted features of the sound to be synthesized to obtain the speech data to be output.

7. An electronic device Characterized in that it includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the speech synthesis model training method according to any one of claims 1 to 5.

8. A computer-readable storage medium Characterized in that the computer-readable storage medium includes a computer program, and when the computer program runs, it controls the electronic device where the computer-readable storage medium is located to execute the speech synthesis model training method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Voice processing method and device, electronic equipment and storage medium

    CN111326136A

  • Phrase-based end-to-end text-to-speech (TTS) synthesis

    CN111681641A